Introducing Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber
Google ● Covered by 19 sources
Google just dropped three new Gemini models: 3.6 Flash, 3.5 Flash-Lite, and a cybersecurity-focused Flash Cyber. They're cheaper, faster, and Gemini 4 training has already started.
Based on reporting by Google — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
Google isn't waiting around. Barely a beat after Gemini 3.5 Flash landed, the company is pushing out 3.6 Flash, 3.5 Flash-Lite, and a specialized model called 3.5 Flash Cyber, all aimed at one target: making AI agents cheap enough and fast enough to run at scale.
The headline number is efficiency. 3.6 Flash uses 17% fewer output tokens than its predecessor on the Artificial Analysis Index, and on some coding benchmarks like DeepSWE, that efficiency gain jumps to 65%. Pair that with a price cut to $1.50 per million input tokens and $7.50 per million output tokens, and Google is essentially betting that the way to win the agent wars is to make every single tool call and reasoning step cost less. The model also posts real gains: 49% versus 37% on DeepSWE, 63.9% versus 49.7% on MLE Bench for ML research work, and a jump to 83.0% on OSWorld-Verified for computer-use tasks. Customers like Hebbia and Harvey are already leaning on it for document parsing and chart analysis.
Then there's 3.5 Flash-Lite, which is less about brains and more about raw speed — 350 output tokens per second, according to Artificial Analysis, at just $0.30 per million input tokens and $2.50 per million output tokens. It's built for the grunt work of agentic pipelines: high-volume search, document processing, the stuff that needs to happen fast and cheap thousands of times a day. Oddly, it now beats the older, bigger 3 Flash on several benchmarks, including SWE-Bench Pro and OSWorld-Verified, which says something about how much low-latency models have caught up to their heavier siblings.
The more unusual release is 3.5 Flash Cyber, a fine-tuned version of Flash built specifically to hunt for and patch software vulnerabilities inside Google's CodeMender agent. Multiple Cyber agents work together to produce a single vulnerability report, and Google claims frontier-competitive results on the CyberGym benchmark. Given that this is explicitly dual-use technology — good for defenders, potentially useful for attackers too — Google is keeping it locked down, offering access only to governments and vetted partners through a limited pilot.
All of this is really scaffolding for what's coming next. Gemini 3.5 Pro is in partner testing now, and Google has quietly confirmed it has already kicked off pretraining for Gemini 4, calling it its most ambitious training run yet. The Flash releases feel less like a destination and more like Google clearing runway for something bigger.
My take — AI-written commentary, not fact-checked reporting
The efficiency-over-raw-power framing here is smart marketing, but it's also just true: nobody running agents at scale cares about a leaderboard, they care about the bill. Locking Flash Cyber behind a government-and-partners pilot is the right call, though it's a reminder that every capability advance in security tooling now comes with an implicit arms-race clause attached. And the fact that Gemini 4 pretraining is already underway while they're still shipping point releases tells you Google isn't slowing down to let anyone catch up.
Read more about this at: Google