Gemini 4 Argon is here: It’s great, and you can’t have it yet
The New Stack Frederic Lardinois ● Covered by 4 sources
Google’s Gemini 4 Argon is out, and it beats OpenAI and Anthropic on most tests. Too bad it’s in a phased rollout, so most people still can’t touch it.
Based on reporting by The New Stack, Frederic Lardinois — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
Google on Wednesday unveiled Gemini 4 Argon, its long-awaited flagship model, and the company is treating it like a carefully staged debut, not a blast-off. The model shows strong results across a wide set of benchmarks, and Google says it is expanding access in phases while it works through the U.S. government’s voluntary pre-release process.
That rollout matters because Argon arrives with the kind of benchmark sheet that invites instant comparison. In Google’s own testing, the model finishes first or shares first place in 13 of 18 benchmarks against OpenAI’s GPT-6 Astra and Anthropic’s Fable 5.1 and Opus 5.5. But the picture is not clean. On some coding tests, Argon leads; on others, it lands last. Google itself points to the split, even while pitching the model as especially strong for office work.
The biggest gap between the hype and the numbers shows up in coding. Google highlights a 77.9% score on DeepSWE v1.1 as a new state of the art, yet Argon trails both GPT-6 Astra and Opus 5.5 on FrontierSWE v2 and Terminal-Bench 4.0. Its 91.9% on Vibe Code Bench looks better, but that field is crowded — all four models are above 89% there. The more interesting story is that Argon seems tuned for practical knowledge work, where it posts 51.3% on Zapier’s AutomationBench, 19.6% on Harvey’s Legal Agent Benchmark, and a strong 84.2% on the longer GraphWalks test.
Google also pushed one notable technical limit: Argon can generate up to one million output tokens, up from 64,000 in earlier Gemini models. That is a big jump, and Google argues it helps the model reason through long problems in a single pass. On the security side, the company says it trained Argon to find, validate, and patch vulnerabilities on its own, and that Fairwind participants and internal teams will get it without cyber guardrails. Pricing starts at $2 per million input tokens and $10 per million output tokens for now, then rises to $4 and $20 later. Paid API customers and AI Ultra subscribers are first in line when access widens, before developers, enterprises, and consumers.
My take — AI-written commentary, not fact-checked reporting
This is the usual AI release theater: a model wrapped in benchmark confetti, then locked behind a phased rollout. The real tell is not the brag sheet, but that Google is betting this one on office work and controlled access, because that’s where the money and the least embarrassing demos live. Everyone else can admire the scores and wait in line like civilized peasants.
Read more about this at: The New Stack