Gemini 3 Deep Think: Advancing science, research and engineering
Google DeepMind
Google just souped up Gemini 3's "Deep Think" mode for hardcore science and engineering work, and it's now testable via API for researchers. It posts wild benchmark scores, but access stays locked behind Ultra subscriptions and an invite-only waitlist.
Based on reporting by Google DeepMind — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
Google DeepMind has pushed out a new version of Gemini 3 Deep Think, its specialized reasoning mode, aimed squarely at the messy, ambiguous problems that show up in real science and engineering rather than tidy textbook exercises. The pitch is straightforward: instead of chasing a single correct answer, the model is tuned to work through incomplete data and open-ended research questions, the kind mathematicians, physicists and engineers actually deal with day to day.
The numbers Google is citing are hard to ignore. Deep Think scored 48.4% on Humanity's Last Exam without tool use, a benchmark specifically built to be brutal for frontier models. On ARC-AGI-2, verified independently by the ARC Prize Foundation, it hit 84.6%. It also posted a Codeforces Elo of 3455, gold-medal level results on the 2025 International Math Olympiad, and matching gold-tier performance on the written portions of this year's Physics and Chemistry Olympiads. On CMT-Benchmark, a test of advanced theoretical physics, it landed at 50.5%. Google says these are the highest marks the company has recorded on several of these evaluations.
But competition math scores don't pay the bills for most labs. So DeepMind is leaning hard on the engineering angle too, showing off a workflow where Deep Think takes a rough hand sketch, works out the underlying 3D geometry, and spits out a file ready for 3D printing. It's a small demo, but it's meant to signal that the model can move from abstract reasoning to something an engineer could actually use on a Tuesday afternoon.
Access is where things get more interesting than the leaderboard numbers. Deep Think has been available inside the Gemini app for Google AI Ultra subscribers, and starting now it's also reachable through the Gemini API — but only for a limited group of researchers, engineers and enterprises who apply through an early access program. Google frames this as a careful rollout in partnership with scientists rather than a mass release, which tracks with how it handled Deep Think's earlier math and coding wins last year.
What's notable is the direction of travel: DeepMind keeps pairing eye-popping benchmark claims with narrower gatekeeping around who actually gets to run the model. That combination — best-in-class scores, drip-fed access — has become a pattern across the frontier lab world, and Gemini 3 Deep Think fits it neatly.
My take — AI-written commentary, not fact-checked reporting
I'll believe Deep Think is genuinely advancing science when it's actually sitting in more labs than DeepMind's PR deck, not gated behind an Ultra subscription and an early-access form. Benchmark gold medals are cheap theater compared to reproducible results from outside researchers, and Google knows it — that's exactly why access stays this tight. Until the tooling is boring and widely available, treat the 84.6% ARC-AGI-2 number as a headline, not a handoff.
Read more about this at: Google DeepMind