Gemini 3.1 Flash-Lite: Built for intelligence at scale
Google DeepMind
Google just launched Gemini 3.1 Flash-Lite, its cheapest and fastest Gemini 3 model yet. It beats last year's Flash model on speed and quality while costing pennies per million tokens.
Based on reporting by Google DeepMind — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
Google DeepMind has slipped a new model into the Gemini 3 lineup, and this one is all about volume. Gemini 3.1 Flash-Lite is now rolling out in preview through AI Studio and Vertex AI, positioned as the model you reach for when you're running millions of requests and every fraction of a cent matters.
The pricing tells the story: $0.25 per million input tokens, $1.50 per million output tokens. That's Flash-Lite territory, but the performance numbers read more like a mid-tier model punching above its weight. According to Artificial Analysis benchmarks, it delivers a Time to First Answer Token 2.5 times faster than Gemini 2.5 Flash, with output speed up 45%. On the Arena.ai leaderboard it posted an Elo of 1432, and it scored 86.9% on GPQA Diamond and 76.8% on MMMU Pro — numbers that, notably, edge out 2.5 Flash, a model from a supposedly higher tier just one generation back.
What's more interesting than the raw scores is the control Google is handing developers. Flash-Lite ships with adjustable thinking levels baked into AI Studio and Vertex AI, letting teams dial reasoning effort up or down depending on the job. Cheap, low-effort inference for bulk translation or content moderation is one mode. Flip the switch and the same model can handle heavier lifting — generating dashboards, running simulations, following multi-step instructions — without needing to swap to a pricier tier.
Early adopters like Latitude, Cartwheel, and Whering are already testing it in production, and Google says feedback points to a model that handles complex inputs with the precision typically reserved for larger, more expensive systems, while sticking closely to instructions. That's the pitch, anyway: less a stripped-down budget option and more a scaling tool that happens to be cheap.
My take — AI-written commentary, not fact-checked reporting
This is Google quietly admitting that most AI workloads don't need a frontier model at all, they need something fast and dirt-cheap that doesn't embarrass itself on reasoning. The adjustable thinking-level trick is the smart part here, letting one model serve as both bulk labor and lightweight reasoner, which is exactly the kind of pragmatic engineering that gets ignored while everyone argues about who has the biggest model.
Read more about this at: Google DeepMind