OpenAI started rolling out GPT-6 Astra by opening access to its next-generation large language model. Astra scored 98% on the FrontierMath Tier 4 test of 50 math challenges, and the rollout had been delayed for several weeks after the model qualified as “critical” for hacking many well-protected systems without human input. OpenAI will now expand Astra to ChatGPT, Codex, and its API over the coming days under the Daybreak program and provide $1 billion in credits for cybersecurity research, training, and support.
OpenAI released GPT-6 Astra as its new flagship model, but first independent benchmarks show it matching its predecessor and trailing top competitors on headline intelligence metrics. Artificial Analysis reports an Intelligence Index score of 61 points, equal to GPT-5.6 Sol and 5 points behind Anthropic’s top model at 66. Costs rise 2.5x with mixed per-token efficiency improvements, while hallucination rate at max effort falls from 92% to 51%.
Simon Willison’s Weblog·22 hours ago·
42
● 12 sources
OpenAI is rolling out GPT-6 Astra to a limited set of organizations before expanding availability to all ChatGPT Plus, Pro, Business, and Enterprise users and to the OpenAI API and AWS. The API pricing is $10 per million input tokens and $50 per million output tokens. Model access expands broadly across ChatGPT and developer platforms, and Astra’s published benchmark results emphasize improvements in ARC-AGI 3 (99.9%) and security/long-context tasks.
OpenAI introduced GPT-6 Astra, positioning it as its most capable model while keeping the frontier version mostly locked behind staged access. Astra is priced at $10 per 1M input tokens and $50 per 1M output tokens via the API. Access expands over coming days from selected Daybreak cybersecurity customers and limited bounded tasks, while the model also includes built-in refusals and additional protections to restrict advanced misuse.
GPT-6 Astra posted a 98.6% result on OpenAI’s ARC-AGI-3 benchmark, far above GPT-5.6 Sol’s 7.8%. The score was evaluated via the Responses API harness with two settings changed for real-world use, and OpenAI notes other models used different setups. The result broadens Astra’s performance claims across other benchmarks and is paired with math progress reports, while OpenAI stresses that the ARC-AGI-3 setup makes direct comparisons less straightforward.
Meta released Muse Spark 1.3, an agentic coding model aimed at long-horizon work with usability features like sustaining long threads and confirming consequential actions. Meta reports it used about 20% fewer tool calls and about 25% fewer tokens than Muse Spark 1.2 in internal comparisons. The model is available today in Muse Code and the Meta Model API (with closed weights), improving coding efficiency and agent run behavior while self-hosting remains unavailable and top reasoning mode stays gated.
Researchers built Spi-Fly, an algorithm inspired by fruit fly smell processing, to better remember scents over time. It was described in a paper published in Neuromorphic Computing and Engineering. The result is an odor-recognition system aimed to reduce the “quick to forget” limitation seen in many commercial electronic noses as it adds longer-lasting scent memory.
A study evaluated how polygenic risk score training strategies transfer from European GWAS data to Japanese samples across eight clinical traits. The largest dataset transferability tests compared training when the target-population (Biobank Japan) sample size reached 15,000 and found that European-only pooling helped only below that cutoff. Model performance then crossed over and increasingly favored target-population-specific training as Biobank Japan sample sizes grew, with conserved traits retaining benefits longer and population-specific traits favoring alternative methods like cross-population meta-analysis.
Researchers, with HHMI Janelia and collaborators, published a complete wiring diagram of the male fruit fly brain and central nervous system, described as the largest brain map to date. The map contains over 166,000 neurons and 125 million synaptic connections. This enables researchers to compare male and female connectomes and use the verified dataset (via Neuroglancer) to study circuit control, behavior, and other neural mechanisms with less manual annotation.
Meta launched Muse Spark 1.3 and said it delivers its biggest improvements yet for coding and agentic tasks, with rollout through Muse Code and the Meta Model API plus open-weight releases planned. It reported 75.4% on the DeepSWE coding benchmark and 98.1% on the 512K-1M MRCR long-context test for Spark 1.3. Outside benchmarks found Spark 1.3 xhigh at 61 on an Intelligence Index at about $0.55 per task, pushing Gemini 3.8 Flash off the cost-versus-intelligence frontier and changing which models look most efficient for developers.
Google DeepMind and Google Research released WeatherNext 3, a new AI weather-forecasting model intended to predict atmospheric behavior more accurately and with faster updates. WeatherNext 3 achieves 5km resolution and is reported to be 60% better at rain than WeatherNext 2 based on evaluations on Operational WeatherBench. Google plans to feed the model into Search, Google Maps, and Gemini, and also offer it to users and researchers via Google Cloud.
Zvi (Don't Worry About the Vase)·1 day ago·
16
● 2 sources
The article surveys a backlog of postmortem coverage after the Hugging Face hack and pivots to a new wave of upcoming or recently mentioned language-model releases. A concrete detail is Claude Code’s weekly limits increase starting September 14, when standard weekly limits rise by 25% for Pro, Max, Team, and seat-based Enterprise plans. As a result, the author plans minimal attention for several claimed step-forward models while focusing coverage next on Fable 5.1 and OpenAI’s Astra, alongside smaller notes on interpretability and usage limits.
OpenAI announced GPT-6 Astra as its next major AI model and said it advances capabilities for tasks like cybersecurity, professional work, software engineering, science, and computer use. The release was described as meeting OpenAI’s critical cybersecurity capability threshold, and it was announced earlier this week. As a result, OpenAI is positioning the model as an AGI-era milestone while also promising it will not be used to replicate prior hacking concerns.
OpenAI launched GPT-6 Astra, its most capable end-to-end model for complex reasoning and software engineering. Pricing is $10 per 1M tokens for short context and $50 per 1M tokens, with initial API model id gpt-6-astra made available today. Availability rolls out first via Trusted Access/Daybreak, then Plus, Pro, Business, Enterprise, and the API in the following days.
NeoMME released a multilingual multimodal encoder family and a NeoMME-Retriever fine-tune for visual document retrieval that uses a single bidirectional Transformer without separate pretrained vision or language towers.
Claude released Fable 5.1, which Ben reports is faster and easier to talk to than the prior model and adds a new system prompt restricting repeated song lyrics and certain copyrighted content. Anthropic also cut the cost of caching inputs for Fable by 75%, making API usage about 25% cheaper than before. As a result, developers can use the model at lower prompt-caching costs while seeing its updated response behavior.
Z.ai announced GLM-5.3 after improving its predecessor GLM-5.2 to raise coding and agentic performance and to gain cybersecurity strength via additional safety testing. GLM-5.3 scored 84.5% on the CyberGym exploit-detection benchmark. Subscriptions and API pricing were set for the rollout, with weights expected about two weeks after launch and a license not yet announced.
Meta rolled out Muse Spark 1.3 and evaluations placed it near the top for coding and agentic performance. Muse Spark 1.3 (xhigh) scored 61 on the Intelligence Index. Its progress mainly improves agentic/scientific reasoning and pricing, and whether open-weight 1.3 arrives could shift the open-weight rankings versus top labs.
Meta’s Muse Spark 1.3 was released, with community claims that it matches GPT-5.6-Sol on multiple evaluations and will ship open weights. The article cites a reported 90%+ pricing discount when users opt in to training. This adds an open-weight option and a much lower training cost for developers building agentic and coding workloads.
Meta released Muse Spark 1.3 as its most powerful large language model so far, aiming to match leading labs like OpenAI and Anthropic. The model scored 62 on Artificial Analysis’s Intelligence Index. Meta will start rolling it out to developers via its Model API and to Facebook, Instagram, and Meta AI in the coming days.
Google launched Gemini 3.8 Flash and Gemini 3.8 Flash Cyber, plus the Fairwind early access program for the cyber-focused model. Gemini 3.8 Flash scored 73.7% on DeepSWE-1.1, outperforming GPT-5.6 Sol by 1% but trailing Claude Opus 5 by a few fractions of a percent. The release shifts Google’s model lineup toward higher-effort reasoning (more iterative tool calls and steps) and expands early access cybersecurity use via CodeMender for selected participants.
LFM2.5-350M was fine-tuned with GRPO using TRL for structured-output compliance and then re-evaluated on IFStruct. The run used about 500 samples and 100 training steps, improving the IFStruct pass rate from 22.6% to 29.7%. JSON compliance increased from 18.0% to 31.9% while YAML stayed roughly flat, indicating task-specific fine-tuning shifts the model toward the targeted output formats.
Every AI story that matters,
in your inbox by 8am.
TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the
day in two minutes. Follow companies and topics for alerts, or get the
briefing in Slack. Free, no spam, unsubscribe anytime.
Reading TLDRocket needs no cookies, and the readership counts we rely on come from
our own cookieless analytics. Google Analytics is the exception: it sets cookies and
reports to Google, so it stays switched off until you allow it. You can change your
mind any time from “Cookie settings” in the footer.
Strictly necessary
Session security and form protection (tldrocket-session,
XSRF-TOKEN, 2 hours). The site cannot work without them,
so they need no consent.
Always on
Google Analytics 4 (_ga,
_ga_<id>, up to 2 years). Measures which
stories and sections readers use. Google acts as a third-party processor and may
store the data outside the EU. No advertising, no profiling, no data sold.