TLDRocket
Sign in

Models & Research

1192 summarised stories in Models & Research, each linking back to the original source. Browse all topics →

Friday, 4 September 2026

ARC Prize test results for Astra on challenging evaluation setups

ARC Prize 8 hours ago 47 5 sources

OpenAI’s GPT-6 Astra tested on ARC-AGI-3 achieved state-of-the-art performance while solving levels by compacting and reusing internal representations. It scored 62.7% on ARC-AGI-3 Semi-Private for $26,098 under the Standard harness and 99.9% for $18,817 under the Provider Adapter harness. Reporting will now include both harness results on the ARC-AGI leaderboard, with Provider Adapter runs also improving speed and reducing total tokens.

Artificial Analysis benchmarks GPT-6 Astra vs other agent models

Artificial Analysis 8 hours ago 16 5 sources

Artificial Analysis reported GPT-6 Astra benchmark results showing it ties Fable 5 on its Coding Agent Index while changing token use and pricing relative to GPT-5.6 Sol. GPT-6 Astra’s input/output token prices were raised to $10/$50 per million from $4/$20 (a 2.5x increase). As a result, Astra is more token efficient and less costly per task in coding, but it is 75% more expensive per task in the Intelligence Index, alongside a drop in hallucination rate from 92% to 51% at max effort and mixed gains across other evaluations.

GPT-6 Astra deep dive: everything you need to know

The Neuron 8 hours ago 22 12 sources

OpenAI introduced GPT-6 Astra as a more tool-using “reasoning model” aimed at getting real work done, backed by demos that involve turning instructions into computer actions and file outputs. Astra’s API model page lists a 1,050,000-token context window. As a result, Astra can handle longer, multi-step projects with tools and memory-like state across requests, and reviewers emphasize that the big ARC-AGI-3 result depends on the specific harness used.

Google DeepMind’s WeatherNext 3 Trains on Weather Station Observations to Deliver 5 km Global Forecasts, Refreshed Every Hour

MarkTechPost 16 hours ago 2 5 sources

Google DeepMind released WeatherNext 3, a global weather forecasting model that uses live geostationary satellite data and reinitializes every hour to improve local detail and reduce latency versus relying on delayed NWP analysis.

Google shipped four Gemini Flash models in 106 days. Yet its Gemini 3.5 Pro is still AWOL.

Fortune 22 6 sources

Google released the Gemini 3.8 Flash AI model while its Gemini 3.5 Pro flagship still has not shipped as expected. Gemini 3.8 Flash passed DeepSWE v1.1 runs with about 74% success at maximum effort, and it cost an average of $2.36 per task versus $11.84 for Anthropic’s Claude Opus 5. Google will continue shipping faster, cheaper Flash updates, but the missing Pro release and internal delays are widening scrutiny over its ability to catch up on the frontier.

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.