The Algorithmic Bridge
·
1 week ago
● 4 sources
OpenAI's GPT-5.6 Sol scored 7.8% on the ARC-AGI-3 benchmark, a test designed to measure fluid intelligence through pattern recognition games that humans solve over 90% of the time. This represents a 20-fold improvement over GPT-5.5's 0.43% score three months earlier, and the model distinguishes itself by correctly identifying game mechanics before execution rather than simply executing learned patterns. The result suggests that further progress toward general intelligence requires improved reasoning scaffolding and planning rather than raw intelligence, since the model's failures occur in multi-step inference composition rather than perception.
Simon Willison
·
1 week ago
● 6 sources
OpenAI released three new GPT-5.6 models—Luna, Terra, and Sol—with input/output token prices of $1/$6, $2.50/$15, and $5/$30 respectively. On Agents' Last Exam benchmark measuring long-running professional workflows across 55 fields, GPT-5.6 Sol scored 53.6 compared to Claude Fable 5's 40.5, while Luna and Terra matched Fable 5's performance at one-sixteenth the cost. The models include new API features for programmatic tool composition, multi-agent spawning, and explicit prompt cache breakpoints, expanding capabilities for agentic workflows.
OpenAI Blog
·
1 week ago
Microsoft has made GPT-5.6 the default model in Microsoft 365 Copilot across its office applications. The model is now integrated into Word, Excel, PowerPoint, Chat, and Cowork. Users of these applications will experience faster processing and improved output quality for AI-assisted tasks.
OpenAI Blog
·
1 week ago
OpenAI launched a bug bounty program focused on biological risks in GPT-5.5, inviting researchers to identify potential misuse cases related to dangerous biotechnology information. Participants can earn up to $2,000 per valid submission for identifying vulnerabilities in the model's safety measures. The program aims to catch biological safety gaps before the model's wider release and integrate researcher findings into OpenAI's safety protocols.
OpenAI Blog
·
1 week ago
● 4 sources
OpenAI released GPT-5.6, a new model that improves efficiency and performance across token usage and computational cost. The model delivers higher capability per dollar spent compared to its predecessors while maintaining quality on complex tasks. Users can now access greater computational power scaled to match the difficulty of their specific workloads.
The Neuron
·
1 week ago
● 6 sources
OpenAI received Commerce Department clearance for GPT-5.6 following additional testing, enabling broad deployment of the model. Three new variants—Sol, Terra, and Luna—launch Thursday. The clearance removes regulatory barriers for OpenAI to compete more directly with Anthropic's offerings.