ARC-AGI-3
Reported model scores on ARC-AGI-3, best parseable score first. Each row keeps its verbatim score, test conditions and provenance, and links to the source coverage it was extracted from. Benchmark profile →
| Model | Score | Conditions | Provenance | Measured | Source |
|---|---|---|---|---|---|
| Claude Opus 5 | 100% | with Nvidia Agentic Variation Operators (AVO) custom agent harness | vendor-reported | — | coverage → |
| Claude Opus 5 | 30.2% | baseline without the Nvidia AVO harness | vendor-reported | — | coverage → |
Scores are only comparable within one benchmark under matching conditions — results under different test setups, and scores from other benchmarks, are not directly comparable.