Google shipped four Gemini Flash models in 106 days. Yet its Gemini 3.5 Pro is still AWOL.
Fortune Wen Shao ● Covered by 4 sources
Google just shipped Gemini 3.8 Flash, its fourth Flash model since May. But the bigger Gemini 3.5 Pro is still missing, and that’s getting awkward.
Based on reporting by Fortune, Wen Shao — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
Google rolled out Gemini 3.8 Flash on Wednesday, and the company is leaning hard on speed, cost, and coding as the selling points. It says the model can match larger rival systems on some benchmarks while running at a much lower cost. The timing is telling: this release lands just three weeks after Gemini 3.7 Flash, and it’s the fourth Flash model Google has shipped since May.
Flash is Google’s name for the small, fast end of Gemini. These models are meant to give users solid answers without burning through as much compute. That part of the story makes sense. What doesn’t is the missing flagship: Gemini 3.5 Pro was supposed to arrive in June, according to Sundar Pichai, but it still hasn’t shown up. Google’s site still calls it “coming soon,” and the Wall Street Journal reported that internal candidates were dropped because they didn’t improve enough over Flash.
That gap is helping fuel a very public bit of shade from rivals. Meta chief AI officer Alexander Wang mocked Google on X after his company’s Muse Spark 1.3 jumped ahead of Google’s models on the Artificial Analysis Intelligence Index. On that ranking, Google’s best model is now Gemini 3.8 Flash, sitting in 10th place. Not exactly the sort of placement that screams “frontier leader.”
Google is trying to turn the Flash cadence into a strength story. The company says the new model was accelerated by long-running AI-agent loops that recursively evaluated and refined the underlying models. Google DeepMind researcher Shunyu Yao called the debut “one small step for model, one giant leap for RSI.” That’s recursive self-improvement territory, or at least something that smells close to it. And for AI researchers, that idea has long been a goal. For AI safety people, it has been one of the scariest ones.
There is still a very practical reason Flash keeps getting the attention. Google said its model APIs were processing about 22 billion tokens per minute in its Q2 earnings call, up from 16 billion the quarter before, while computing supply remained constrained. Flash is the workhorse line, Pichai said, because it hits the sweet spot between performance and cost. So Google can keep pushing the cheaper models, even if the bigger Pro release is stuck in the hallway.
On some coding tests, the new Flash model is genuinely competitive. In DeepSWE v1.1, Gemini 3.8 Flash at high effort and Anthropic’s Claude Opus 5 at maximum effort each passed about 74% of scored runs, though Gemini’s average task cost was far lower at $2.36 versus $11.84. Google kept the introductory API price at 75 cents per million input tokens and $3.75 per million output tokens. Even so, Artificial Analysis said the new model cost about 40% more per task than Gemini 3.7 Flash, because it used more output tokens and took more agent steps. Faster doesn’t always mean cheaper. And in this case, it also doesn’t mean the flagship is ready.
My take — AI-written commentary, not fact-checked reporting
Google’s Flash obsession looks less like strategic genius than a company making the best of a missing headline act. Shipping faster small models is useful; calling it destiny for RSI is classic AI-lab self-congratulation with better marketing. The real test is still the one Google keeps postponing: where’s Pro?
Read more about this at: Fortune
Related stories
Google announces Gemini 3.6 Flash and cybersecurity AI, teases 3.5 Pro and Gemini 4
Ars Technica · 1 month ago ·
39