TLDRocket
Sign in

[AINews] Gemini 4 Argon: GDM’s answer to Astra/Fable, with 1M output

Latent Space ● Covered by 11 sources

Google DeepMind dropped Gemini 4 Argon, but only for trusted cyber testers for now. It claims a 1M-token output limit and strong benchmark wins, which is a pretty loud comeback.

Based on reporting by Latent Space — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Google DeepMind has finally answered the Astra and Fable wave with Gemini 4 Argon, a new model aimed at coding, enterprise work, and cyber defense. That alone would be enough to get attention. The bigger twist is the rollout: access starts in the Fairwind Program with government users and trusted cyber defenders, while Google says broader access to developers, enterprises, and consumers will come later after more guardrail work.

Argon arrives after a long stretch where Google’s biggest model news was more incremental. GDM last shipped a model larger than Flash in February with 3.1 Pro, then kept pushing 3.x Flash versions while the company went through a management shakeup last month. So this is less a routine launch than a statement that Google wants back in the front line.

On paper, Argon looks strong. Google says it is first on 13 of 19 published benchmarks against GPT-6 Astra and Claude Opus 5.5. The standout claim is output length: Google cites an industry-leading 1M-token output limit, up from 64K. Artificial Analysis says that ceiling is reachable through Long Decode Continuation, a new API feature that pauses a long response and resumes it across calls. Vals, meanwhile, lists a 262K maximum output, so the exact practical number depends on how you measure it.

The pricing is straightforward, if not exactly cheap. Standard rates are $4 and $20 per 1M input and output tokens, with a 50% introductory discount cutting that to $2 and $10. Cached input gets a 95% discount. Artificial Analysis also says Argon matches GPT-6 Astra on its Intelligence Index at 53, but does so with different economics: cheaper per task at the discounted rate, yet using more output tokens on average than Astra.

There’s a lot more in the eval pile. Google says Argon helped free more than 300 TiB of data-center memory and is being used to migrate more than 800K lines of C/C++ kernel code to Rust. External evaluations put it near the top on AutomationBench-AA, Vals Index, Vibe Code Bench, and several terminal and security tests, while also showing some weaker spots like a legal benchmark where it trails Muse Spark 1.2. And yes, people are already arguing about the numbers, including complaints about preference-data benchmaxxing and skepticism around some of the published figures. That seems about right for a frontier launch in 2026.

My take — AI-written commentary, not fact-checked reporting

Google loves a controlled rollout when the model can still scare the room. A cyber-first preview for a 1M-token system is sensible, but it also says the quiet part out loud: frontier AI is now being handed to defenders first because everyone expects the offensive side to show up five minutes later. The benchmark bragging is nice; the real test is whether Argon stays impressive once it leaves the velvet rope and meets actual users.

Read more about this at: Latent Space

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.