ARC-AGI-3
Benchmark ● Covered in 11 stories + Follow
ARC-AGI-3 is an AI benchmark featured across recent coverage of GPT-6 Astra and other agent systems. Recent reports show GPT-6 Astra reaching widely varying ARC-AGI-3 scores depending on the evaluation setup—e.g., 62.7% vs 99.9% across different harnesses—and parallel findings that changes to “harness” and agent wrapper components can materially affect results. The benchmark has also been used to highlight ongoing debate over how much of measured performance comes from model capability versus evaluation and agent-system design.
Updated 17 September 2026
Latest developments
2026
OpenAI launches GPT-6 Astra and begins a staged rollout to selected users and partners, followed by broader availability Model release
OpenAI launches GPT-6 Astra, a computer-use model with a 1.05M-token context window and staged enterprise/API access Model release
Nvidia reported that a custom agent harness using Agentic Variation Operators enabled Claude Opus 5 to achieve a 100% score on the ARC-AGI-3 benchmark Benchmark result
Prime Intellect Open-Sources Prime Agent, an RLM-based Coding Agent Using Persistent IPython Kernel Open source release
- GPT-6 Astra, Looped Transformers, and Hidden Reasoning
- GPT-6: How AGI Became a Case for the Marketing Department
- 🔮 Astra outruns visibility EV#600
- OpenAI will sell you Astra, but not the system that scored 98.6% on ARC-AGI-3
- ARC Prize test results for Astra on challenging evaluation setups
- GPT-6 Astra deep dive: everything you need to know
- GPT-6 Astra: an automated AI Engineer you can hire for
- GPT-6 Astra aced the hardest AI benchmark. The asterisk matters more than the score.
- Nvidia just showed that the harness, not the AI model, is now the real hero
- Claude Opus 5 scored 30% on ARC-AGI-3. Wrapped in Nvidia’s AVO, it hit 100%.
- Prime Intellect Releases Prime Agent: An Open-Source RLM Harness Where Sub-Agents Are Function Calls Inside Persistent IPython Kernel
Relationships
Products & technology
- GPT-6 Astra integrated with this benchmark · 2 sources
- NVIDIA deploys this benchmark · 2 sources
- OpenAI develops this benchmark · 1 source
- GPT-6 Astra deploys this benchmark · 1 source
- GPT-5.6 Sol deploys this benchmark · 1 source
- OpenAI deploys this benchmark · 1 source
- ARC Prize integrated with this benchmark · 1 source
- Claude Opus 5 deploys this benchmark · 1 source