TLDRocket
Sign in

Google AI Just Released Gemini 3.7 Flash: A Coding and Agent Model at $0.75/1M Input Tokens

MarkTechPost Asif Razzaq Covered by 7 sources

Google launched Gemini 3.7 Flash, a cheaper update to 3.6 Flash. It’s aimed at coding and agents, and it’s half the old list price.

Based on reporting by MarkTechPost, Asif Razzaq — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Google has pushed out Gemini 3.7 Flash, only three weeks after 3.6 Flash. This one isn’t a fresh pretraining run. Google says it’s a refinement, built from algorithmic changes to the model’s reasoning core. That framing matters: this is an upgrade meant to squeeze more out of the same basic machine, not a grand reset.

The pricing is the loudest part of the launch. Gemini 3.7 Flash comes in at $0.75 per 1M input tokens and $3.75 per 1M output tokens. That is half the original 3.6 Flash list rate, and Google also pitches it as about a third of the blended cost of Claude Sonnet 5 or GPT-5.6 Terra. For teams trying to keep agent costs from spiraling, that kind of cut is not subtle.

And the model is built for exactly the places where cheap, persistent inference matters most. Google says the gains land in software engineering, document-heavy knowledge work, and web development. The model takes text, images, audio, and video, runs across a 1M-token context window, and can return up to 64K output tokens. It also supports customizable thinking settings, so users can trade quality against cost and latency.

The benchmark picture is strong in some spots and less tidy in others. On FrontierCode 1.1 Main, it scores 43.6% versus 34.4% for 3.6 Flash. On DeepSWE v1.1 it reaches 65.3%, and on WebDev Arena it posts an Elo of 1588, which Google says is the top number in its comparison table. Document tasks move too: GDP.pdf rises from 22.0% to 34.0%, and AutomationBench jumps from 17.0% to 30.4%, ahead of both Claude Sonnet 5 and GPT-5.6 Terra in that set.

But the picture isn’t pure upside. Google’s own tables show GPT-5.6 Terra ahead on several evals, including DeepSWE, Terminal-bench 2.1, Terminal-bench 3.0, and OSWorld-2.0. CharXiv Reasoning also dips a bit from 85.2% to 84.5% without tools. So this is not a clean “best model everywhere” release. It’s a targeted cost-and-capability play, with the clearest wins in agents, PDFs, and code-heavy work.

My take — AI-written commentary, not fact-checked reporting

This is the part of the AI market that actually makes sense: cheaper models that can stay on all day and do boring work without burning money like a small bonfire. The catch, of course, is the usual Google tax — hosted-only, no open weights, no self-hosting, no escape hatch for the people who need one. If the pitch is “build agents at scale,” then locking out air-gapped and residency-heavy teams is a very specific kind of self-own.

Read more about this at: MarkTechPost

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.