Google DeepMind releases Gemini 2.5 Deep Think with gold-medal performance on International Mathematical Olympiad and Programming Contest benchmarks
Model release ● Confirmed 92% confidence first seen
Google DeepMind released Deep Think, a reasoning model variant of Gemini, which achieved gold-medal level performance on the 2025 International Mathematical Olympiad by solving 5 of 6 problems (35/42 points) and on the 2025 International Collegiate Programming Contest World Finals by solving 10 of 12 problems. The model is now available in the Gemini app for AI Ultra subscribers with limitations on daily usage.
Decision brief
- What changed
- Google DeepMind released Gemini 2.5 Deep Think, a reasoning-focused model that achieved officially certified gold-medal performance at the 2025 International Mathematical Olympiad (35/42 points, 5 of 6 problems, natural language only) and gold-medal level at the ICPC World Finals (10 of 12 problems). The version available in the Gemini app for AI Ultra subscribers is a faster but lower-capability variant reaching only bronze-level IMO performance, with daily usage limits.
- Why it matters
- This signals rapid progress in AI's ability to handle multi-step abstract reasoning and competitive-level problem solving, which has direct implications for software engineering, R&D, and technical hiring pipelines. However, the gap between the gold-medal competition model and the bronze-level consumer product means near-term practical capability for most users is more limited than headlines suggest, so leaders should calibrate expectations and pilot scope accordingly.
- Evidence
- All three articles are Google DeepMind's own releases/blog posts describing the same event, so coverage is consistent but not independently verified by third-party press; IMO scoring is described as officially certified by IMO coordinators, adding some external validation for that specific benchmark.
- What remains uncertain
- It's unclear how the bronze-level consumer product performance translates to real-world business tasks versus curated competition problems, and no independent or third-party analysis is included to corroborate DeepMind's self-reported results. The specific 'fixed number of prompts per day' limit and broader availability timeline beyond AI Ultra subscribers are also unspecified.
- Monitor next
- Watch for independent benchmarking or enterprise case studies testing Deep Think on real-world coding/research tasks outside controlled competition settings.
Analytical support, not advice — assumptions and open questions stated above.