TLDRocket
Sign in

Models & Research

1192 summarised stories in Models & Research, each linking back to the original source. Browse all topics →

Thursday, 3 September 2026

OpenAI starts rolling out its next-generation GPT-6 Astra model

SiliconANGLE 19 hours ago 46 12 sources

OpenAI started rolling out GPT-6 Astra by opening access to its next-generation large language model. Astra scored 98% on the FrontierMath Tier 4 test of 50 math challenges, and the rollout had been delayed for several weeks after the model qualified as “critical” for hacking many well-protected systems without human input. OpenAI will now expand Astra to ChatGPT, Codex, and its API over the coming days under the Daybreak program and provide $1 billion in credits for cybersecurity research, training, and support.

GPT-6 Astra Trails Top Models From Anthropic and Meta in Benchmarks

Trending Topics 21 hours ago 24 5 sources

OpenAI released GPT-6 Astra as its new flagship model, but first independent benchmarks show it matching its predecessor and trailing top competitors on headline intelligence metrics. Artificial Analysis reports an Intelligence Index score of 61 points, equal to GPT-5.6 Sol and 5 points behind Anthropic’s top model at 66. Costs rise 2.5x with mixed per-token efficiency improvements, while hallucination rate at max effort falls from 92% to 51%.

GPT‑6 Astra

Simon Willison’s Weblog 22 hours ago 42 12 sources

OpenAI is rolling out GPT-6 Astra to a limited set of organizations before expanding availability to all ChatGPT Plus, Pro, Business, and Enterprise users and to the OpenAI API and AWS. The API pricing is $10 per million input tokens and $50 per million output tokens. Model access expands broadly across ChatGPT and developer platforms, and Astra’s published benchmark results emphasize improvements in ARC-AGI 3 (99.9%) and security/long-context tasks.

GPT-6, Also Known as “Astra,” Is Here to Beat Anthropic and Be “AGI”

Trending Topics 23 hours ago 1 5 sources

OpenAI introduced GPT-6 Astra, positioning it as its most capable model while keeping the frontier version mostly locked behind staged access. Astra is priced at $10 per 1M input tokens and $50 per 1M output tokens via the API. Access expands over coming days from selected Daybreak cybersecurity customers and limited bounded tasks, while the model also includes built-in refusals and additional protections to restrict advanced misuse.

GPT-6 Astra aced the hardest AI benchmark. The asterisk matters more than the score.

The New Stack 23 hours ago 33 5 sources

GPT-6 Astra posted a 98.6% result on OpenAI’s ARC-AGI-3 benchmark, far above GPT-5.6 Sol’s 7.8%. The score was evaluated via the Responses API harness with two settings changed for real-world use, and OpenAI notes other models used different setups. The result broadens Astra’s performance claims across other benchmarks and is paired with math progress reports, while OpenAI stresses that the ARC-AGI-3 setup makes direct comparisons less straightforward.

Meta AI Released Muse Spark 1.3: An Agentic Coding Model That Uses ~20% Fewer Tool Calls and ~25% Fewer Tokens Than Muse Spark 1.2

MarkTechPost 23 hours ago 47 8 sources

Meta released Muse Spark 1.3, an agentic coding model aimed at long-horizon work with usability features like sustaining long threads and confirming consequential actions. Meta reports it used about 20% fewer tool calls and about 25% fewer tokens than Muse Spark 1.2 in internal comparisons. The model is available today in Muse Code and the Meta Model API (with closed weights), improving coding efficiency and agent run behavior while self-hosting remains unavailable and top reasoning mode stays gated.

Just like a fruit fly, a new algorithm never forgets old scents

Ars Technica 23 hours ago 39

Researchers built Spi-Fly, an algorithm inspired by fruit fly smell processing, to better remember scents over time. It was described in a paper published in Neuromorphic Computing and Engineering. The result is an odor-recognition system aimed to reduce the “quick to forget” limitation seen in many commercial electronic noses as it adds longer-lasting scent memory.

Transfer learning for genomic prediction in underrepresented populations

Google Research 1 day ago 7

A study evaluated how polygenic risk score training strategies transfer from European GWAS data to Japanese samples across eight clinical traits. The largest dataset transferability tests compared training when the target-population (Biobank Japan) sample size reached 15,000 and found that European-only pooling helped only below that cutoff. Model performance then crossed over and increasingly favored target-population-specific training as Biobank Japan sample sizes grew, with conserved traits retaining benefits longer and population-specific traits favoring alternative methods like cross-population meta-analysis.

A connectomics milestone: Mapping the complete male fruit fly brain

Google Research 1 day ago 25

Researchers, with HHMI Janelia and collaborators, published a complete wiring diagram of the male fruit fly brain and central nervous system, described as the largest brain map to date. The map contains over 166,000 neurons and 125 million synaptic connections. This enables researchers to compare male and female connectomes and use the verified dataset (via Neuroglancer) to study circuit control, behavior, and other neural mechanisms with less manual annotation.

“Google was ahead only a few hours”: Muse Spark 1.3 edges out Gemini as Meta claims its biggest coding leap yet

The New Stack 1 day ago 25 8 sources

Meta launched Muse Spark 1.3 and said it delivers its biggest improvements yet for coding and agentic tasks, with rollout through Muse Code and the Meta Model API plus open-weight releases planned. It reported 75.4% on the DeepSWE coding benchmark and 98.1% on the 512K-1M MRCR long-context test for Spark 1.3. Outside benchmarks found Spark 1.3 xhigh at 61 on an Intelligence Index at about $0.55 per task, pushing Gemini 3.8 Flash off the cost-versus-intelligence frontier and changing which models look most efficient for developers.

Google’s latest AI weather model gives you no excuse to forget your umbrella

TechCrunch 1 day ago 43 5 sources

Google DeepMind and Google Research released WeatherNext 3, a new AI weather-forecasting model intended to predict atmospheric behavior more accurately and with faster updates. WeatherNext 3 achieves 5km resolution and is reported to be 60% better at rain than WeatherNext 2 based on evaluations on Operational WeatherBench. Google plans to feed the model into Search, Google Maps, and Gemini, and also offer it to users and researchers via Google Cloud.

AI #184: Post Post Mortem

Zvi (Don't Worry About the Vase) 1 day ago 16 2 sources

The article surveys a backlog of postmortem coverage after the Hugging Face hack and pivots to a new wave of upcoming or recently mentioned language-model releases. A concrete detail is Claude Code’s weekly limits increase starting September 14, when standard weekly limits rise by 25% for Pro, Max, Team, and seat-based Enterprise plans. As a result, the author plans minimal attention for several claimed step-forward models while focusing coverage next on Fable 5.1 and OpenAI’s Astra, alongside smaller notes on interpretability and usage limits.

OpenAI’s next big AI model has ‘entered the AGI era’

The Verge 1 day ago 19 12 sources

OpenAI announced GPT-6 Astra as its next major AI model and said it advances capabilities for tasks like cybersecurity, professional work, software engineering, science, and computer use. The release was described as meeting OpenAI’s critical cybersecurity capability threshold, and it was announced earlier this week. As a result, OpenAI is positioning the model as an AGI-era milestone while also promising it will not be used to replicate prior hacking concerns.

GPT-6 Astra

Product Hunt 1 day ago 35 12 sources

OpenAI launched GPT-6 Astra, its most capable end-to-end model for complex reasoning and software engineering. Pricing is $10 per 1M tokens for short context and $50 per 1M tokens, with initial API model id gpt-6-astra made available today. Availability rolls out first via Trusted Access/Daybreak, then Plus, Pro, Business, Enterprise, and the API in the following days.

Fable 5.1

Ben's Bites 1 day ago 15 3 sources

Claude released Fable 5.1, which Ben reports is faster and easier to talk to than the prior model and adds a new system prompt restricting repeated song lyrics and certain copyrighted content. Anthropic also cut the cost of caching inputs for Fable by 75%, making API usage about 25% cheaper than before. As a result, developers can use the model at lower prompt-caching costs while seeing its updated response behavior.

GLM-5.3’s Exploits, AI Models and Hardware Speed Up, DeepSeek’s New Agent Harness

The Batch 41

Z.ai announced GLM-5.3 after improving its predecessor GLM-5.2 to raise coding and agentic performance and to gain cybersecurity strength via additional safety testing. GLM-5.3 scored 84.5% on the CyberGym exploit-detection benchmark. Subscriptions and API pricing were set for the rollout, with weights expected about two weeks after launch and a license not yet announced.

Muse Spark 1.3: Meta Is Back at the Top, and the Best Open-Weight Model Could Follow

Trending Topics 1 day ago 7 2 sources

Meta rolled out Muse Spark 1.3 and evaluations placed it near the top for coding and agentic performance. Muse Spark 1.3 (xhigh) scored 61 on the Intelligence Index. Its progress mainly improves agentic/scientific reasoning and pricing, and whether open-weight 1.3 arrives could shift the open-weight rankings versus top labs.

[AINews] Muse Spark 1.3 matches GPT-5.6-Sol, confirming Meta Superintelligence as the newest Frontier Lab, >90% discount for training

Latent Space 1 day ago 9 8 sources

Meta’s Muse Spark 1.3 was released, with community claims that it matches GPT-5.6-Sol on multiple evaluations and will ship open weights. The article cites a reported 90%+ pricing discount when users opt in to training. This adds an open-weight option and a much lower training cost for developers building agentic and coding workloads.

Meta says it has caught up with Anthropic and OpenAI with Muse Spark 1.3, its most powerful AI model yet

SiliconANGLE 1 day ago 2 2 sources

Meta released Muse Spark 1.3 as its most powerful large language model so far, aiming to match leading labs like OpenAI and Anthropic. The model scored 62 on Artificial Analysis’s Intelligence Index. Meta will start rolling it out to developers via its Model API and to Facebook, Instagram, and Meta AI in the coming days.

Google launches two Gemini 3.8 models with cutting-edge reasoning capabilities

SiliconANGLE 1 day ago 12 3 sources

Google launched Gemini 3.8 Flash and Gemini 3.8 Flash Cyber, plus the Fairwind early access program for the cyber-focused model. Gemini 3.8 Flash scored 73.7% on DeepSWE-1.1, outperforming GPT-5.6 Sol by 1% but trailing Claude Opus 5 by a few fractions of a percent. The release shifts Google’s model lineup toward higher-effort reasoning (more iterative tool calls and steps) and expands early access cybersecurity use via CodeMender for selected participants.

Fine-tuning a 350M Model for Better Structured Outputs in 100 GRPO Steps

Hugging Face 1 day ago 33

LFM2.5-350M was fine-tuned with GRPO using TRL for structured-output compliance and then re-evaluated on IFStruct. The run used about 500 samples and 100 training steps, improving the IFStruct pass rate from 22.6% to 29.7%. JSON compliance increased from 18.0% to 31.9% while YAML stayed roughly flat, indicating task-specific fine-tuning shifts the model toward the targeted output formats.

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.