Introducing Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber
Google DeepMind ● Covered by 19 sources
Google dropped three new Gemini models: 3.6 Flash, 3.5 Flash-Lite, and a cybersecurity-focused Flash Cyber for CodeMender. They're cheaper, faster, and burn way fewer tokens — while Gemini 4 training has already begun.
Based on reporting by Google DeepMind — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
Google isn't waiting around. Barely any time after shipping Gemini 3.5 Flash, DeepMind is back with a trio of updates aimed squarely at the unglamorous but expensive part of AI: running agents at scale without torching your token budget.
The headline act is 3.6 Flash, pitched as the new workhorse for coding and knowledge work. Google says it uses 17% fewer output tokens than 3.5 Flash on the Artificial Analysis Index, and on DeepSWE, a coding benchmark from Datacurve, the savings hit as high as 65%. It also gets measurably better at the actual work — 49% versus 37% on DeepSWE precision, 63.9% versus 49.7% on MLE Bench for ML research tasks, and a jump to 83.0% on OSWorld-Verified for computer-use tasks. Pricing lands at $1.50 per million input tokens and $7.50 per million output tokens, undercutting its predecessor while doing more per dollar. Customers like Hebbia and Harvey are apparently already leaning on it for document parsing and report drafting.
Then there's 3.5 Flash-Lite, which Google is calling the fastest model in the 3.5 lineup at 350 output tokens per second, priced at a mere $0.30 per million input tokens and $2.50 per million output tokens. It's built for the grunt work — agentic search, document processing, high-volume production traffic — and developers can dial its thinking level up or down depending on whether they need speed or depth. Oddly enough, on several benchmarks it beats the larger 3 Flash model, including SWE-Bench Pro (54.2% vs 49.6%) and OSWorld-Verified (74.0% vs 65.1%). That's a small model punching above its weight class, and it's the kind of efficiency gain that actually shows up on a cloud bill.
The more unusual release is 3.5 Flash Cyber, a specialized variant fine-tuned for hunting and patching software vulnerabilities, deployed through Google's CodeMender agent. Multiple Flash Cyber instances work together to produce a single vulnerability report, and Google claims it's competitive with frontier models on the CyberGym benchmark despite running on a much cheaper base model. Because a tool that finds security holes fast can just as easily be used to exploit them, Google is keeping this one on a short leash — limited to governments and vetted partners through a pilot program rather than a public API release.
Buried near the bottom of the announcement is the bigger story: Gemini 3.5 Pro is already in testing with partners, and Google says it has kicked off its most ambitious pre-training run yet for Gemini 4. Flash models get the press release, but the frontier race clearly hasn't slowed down.
My take — AI-written commentary, not fact-checked reporting
Google burying
Read more about this at: Google DeepMind