Google releases Gemini 3 AI model with multimodal capabilities and agentic features
Model release ● Confirmed 95% confidence first seen
Google released Gemini 3, its latest AI model featuring multimodal understanding, reasoning, and autonomous agent capabilities. The model achieved notable benchmark scores including 1501 Elo on LMArena Leaderboard and 91.9% on GPQA Diamond, and is now available across Google Search, the Gemini app, developer platforms including the new Google Antigravity, and Vertex AI for enterprises. Gemini 3 Pro is designed for developer assistance with coding and app creation, priced at $2 per million input tokens through the Gemini API.
Decision brief
- What changed
- Google released Gemini 3, a new multimodal AI model with reasoning and autonomous agentic capabilities, deployed across Search, the Gemini app, developer tools (including the new Google Antigravity platform), and Vertex AI for enterprises. Gemini 3 Pro is priced at $2 per million input tokens and posted strong benchmark results (1501 Elo on LMArena, 91.9% on GPQA Diamond, 54.2% on Terminal-Bench 2.0).
- Why it matters
- This is a direct competitive move in the frontier AI model race, with immediate enterprise access via Vertex AI and integration into widely used developer tools like Cursor, GitHub Copilot, and Android Studio, lowering adoption friction. Independent testing (One Useful Thing) suggests the model can autonomously complete complex, multi-step research and coding tasks with minimal human guidance, signaling a shift toward agentic AI that could change how technical and knowledge work is scoped and staffed. Leaders evaluating AI vendor strategy, build-vs-buy decisions, and workforce planning for technical roles should treat this as a signal of accelerating capability and lower cost per task.
- Evidence
- Coverage comes from Google DeepMind's own announcements (product and pricing details) plus one independent hands-on account (One Useful Thing) describing autonomous completion of a research paper; benchmark figures are self-reported by Google and not independently verified in this coverage.
- What remains uncertain
- Benchmark scores (Elo, GPQA, ARC-AGI-2, Terminal-Bench) are vendor-reported and not corroborated by third-party evaluation in this coverage; real-world reliability, error rates, and safety/oversight requirements for agentic tasks remain unverified. It's also unclear how pricing and performance compare directly to competing frontier models (e.g., OpenAI, Anthropic) since no head-to-head coverage is provided.
- Monitor next
- Watch for independent third-party benchmarking or enterprise case studies validating Gemini 3's agentic task performance and reliability outside Google's own testing and promotional coverage.
Analytical support, not advice — assumptions and open questions stated above.