A new era of intelligence with Gemini 3
Google DeepMind ● Covered by 3 sources
Google just launched Gemini 3, and it's rolling out everywhere at once—Search, the Gemini app, and dev tools, all today. It tops the leaderboards on reasoning and coding benchmarks, which is Google's boldest AI claim yet.
Based on reporting by Google DeepMind — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
Google didn't ease into this one. Gemini 3 Pro landed simultaneously in Search's AI Mode, the Gemini app, AI Studio, Vertex AI, and a brand-new coding platform called Antigravity. Sundar Pichai called it the first time Gemini has shipped in Search on day one, which tells you Google wants this rollout to feel less like a research preview and more like a product launch with the full weight of the company behind it.
The benchmark numbers are the kind Google clearly wants people to sit with. Gemini 3 Pro hit 1501 Elo on LMArena, scored 37.5% on Humanity's Last Exam without tool use, and posted 91.9% on GPQA Diamond. On coding specifically, it topped WebDev Arena at 1487 Elo and jumped to 76.2% on SWE-bench Verified, a sizable leap from 2.5 Pro. Google is also teasing a Deep Think mode that pushes further still, hitting 45.1% on ARC-AGI-2, a benchmark built to test genuinely novel problem-solving rather than pattern matching. That one's not public yet — it's going to safety testers first before Ultra subscribers get it.
What's more interesting than the scores is the framing around personality. Google says Gemini 3 is built to be less sycophantic, trading flattery for what it calls genuine insight — telling users what they need to hear, not what they want to hear. That's a direct jab at a criticism that's dogged chatbots for the past year: models that agree with everything and praise mediocre ideas to keep users happy. Whether Gemini 3 actually pulls that off in practice, rather than in a benchmark writeup, is the kind of thing only real usage will settle.
Then there's Antigravity, the new agentic development platform Google is positioning as a genuine departure from the usual AI-assisted IDE. Instead of an assistant bolted onto an editor, agents get direct access to the editor, terminal, and browser, letting them plan and execute full software tasks and check their own work through actual browser interaction. Paired with Gemini 3's long-horizon planning — Google points to a Vending-Bench 2 test where the model ran a simulated vending machine business for a full year without drifting off-task — the pitch is that Gemini is inching from chatbot toward autonomous operator, at least for narrow, bounded jobs like booking services or managing an inbox.
My take — AI-written commentary, not fact-checked reporting
I'll believe the sycophancy fix when I see it survive actual users pushing back on bad prompts, not just a benchmark writeup. The bigger story here is Google shipping simultaneously across Search, an app with 650 million monthly users, and enterprise tools — that's the real moat, not the Elo score, and it's the part rivals without Google's distribution simply can't match.
Read more about this at: Google DeepMind