TLDRocket
Sign in

Google Finally Brings Gemini 4 Into the Fight Against Anthropic and OpenAI

Trending Topics Jakob Steinschaden ● Covered by 7 sources

Google unveiled Gemini 4 Argon, its first big model in nearly a year. It says Argon beats Anthropic and OpenAI in many tests, but nobody outside Google can check yet.

Based on reporting by Trending Topics, Jakob Steinschaden — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Google says it’s back in the front row. The company has launched Gemini 4 Argon, its first real flagship model in almost a year, and on Google’s own numbers it tops Anthropic and OpenAI in most of the benchmarks it chose to show off.

The bigger technical move is the size of the response window. Argon can produce up to one million tokens in a single answer, up from 64,000 before. Google’s pitch is simple: let the model tackle a hard problem in one long stretch instead of forcing it through a chain of smaller steps. That fits the use cases Google keeps repeating — software engineering, legal and finance work, and cybersecurity defense.

Google is already using Argon internally, according to its blog post. It says thousands of employees use it for coding, research and writing. The examples it cites are very Google: memory optimizations in its data centers that freed up more than 300 tebibytes, and a migration of large C and C++ codebases to Rust. It also says a video decoder optimized by Argon agents now runs 2.7 times as fast as the previous Rust version.

Pricing is part of the pitch too. Argon is set to start at $2 per million input tokens and $10 per million output tokens, which Google says is half of what Anthropic charges for Claude Opus 5.5. In Google’s own table of 19 benchmarks, Argon finishes first outright 13 times and ties for first once. It leads on the Vals Index at 68.9 percent, ahead of Claude Opus 5.5, Claude Fable 5.1 and GPT-6 Astra. It also posts 51.3 percent on Zapier’s AutomationBench and 77.9 percent on DeepSWE v1.1.

But the table has asterisks everywhere, even if Google doesn’t print them that way. Claude Opus 5.5 still wins on Terminal-Bench 4.0 and PostTrainBench, while GPT-6 Astra takes FrontierSWE v2, Terminal-Bench Science and OSWorld-2.0. And none of this has been checked independently, because Argon is not publicly available yet. Even inside Google, according to Bloomberg’s reporting, there are doubts that the model works as cleanly in real use as it does on benchmarks. That’s the familiar AI story: the scoreboard looks great right up until people have to use the thing.

My take — AI-written commentary, not fact-checked reporting

Google is doing the classic benchmark victory lap before the public can poke holes in it, which is exactly why these launches should be treated like teaser trailers, not verdicts. The more interesting signal is that the strongest version is going to cybersecurity teams first; everyone says “safety,” but the real order of operations is usually “useful first, trustworthy later.” That’s the industry’s favorite magic trick.

Read more about this at: Trending Topics

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.