Cyber and biological tasks
Reported model scores on Cyber and biological tasks, best parseable score first. Each row keeps its verbatim score, test conditions and provenance, and links to the source coverage it was extracted from. Benchmark profile →
| Model | Score | Conditions | Provenance | Measured | Source |
|---|---|---|---|---|---|
| Claude Opus 4.7 | Consistently refused offensive requests | SaferAI testing | independent | — | coverage → |
| GLM-5.2 | Refused none of the offensive tasks | SaferAI testing | independent | — | coverage → |
Scores are only comparable within one benchmark under matching conditions — results under different test setups, and scores from other benchmarks, are not directly comparable.