TLDRocket
Sign in

One Model Family, Two Gold-Level Results: Fine-Tuning Nemotron for IOI and IMO

Hugging Face

Nemotron was tuned into gold-level solvers for both coding and math contests. The catch: it needed training plus a search-and-check loop, not just a bigger model.

Based on reporting by Hugging Face — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Nemotron just pulled off a neat trick: one model family was pushed into gold-medal territory in two very different olympiads. At the International Olympiad in Informatics, the specialized system scored 535.4 out of 600, above the gold threshold and even above the top human score cited in the post. At the International Mathematical Olympiad, the system scored 30 out of 42, clearing the gold line of 29.

The two contests reward almost opposite kinds of intelligence. IOI is about algorithms, code, hidden tests, and time pressure. IMO is about natural-language proofs that can survive a human grader. Getting one model family to work in both settings is the real point here, because it suggests the lesson is not “build a better foundation model for every task,” but “specialize the right foundation model well.”

For coding, the team started from Nemotron 3 and built two specialist versions from a curated set of 22,000 problems plus synthetic reasoning traces. Nemotron-3-Nano-CC, with 30 billion total parameters and 3 billion active parameters, used supervised fine-tuning and reinforcement learning. Nemotron-3-Ultra-CC, with 550 billion total parameters and 55 billion active parameters, used supervised fine-tuning. The gains were large: Nano moved from 130 points before post-training to 280 after SFT and 291 after RL, then to 468 with GenCorrect, the generate-evaluate-refine loop. Ultra-CC reached 502 with the same test-time strategy, and later 535.4 in the IOI 2026 run.

The math side followed the same philosophy, but with proof writing instead of code. Starting from Nemotron 3 Ultra, one specialist was trained with supervised fine-tuning on 414,890 filtered examples across 15,818 unique proof problems, while another was trained with RL on 9,597 problems chosen near the model’s capability frontier. The final IMO system used both, plus the general model, to generate proofs, critique them, and refine the best attempts before a separate high-compute stage picked the submission. It worked in natural language only, with no formal prover, no external tools, and no internet access.

The bigger takeaway is that the medals came from a pair of systems, not a single magic checkpoint. Better fine-tuning made the search loop smarter, and the search loop made the fine-tuning matter more. That’s the boring truth, which is usually the useful one.

My take — AI-written commentary, not fact-checked reporting

This is the open-model version of the obvious lesson Silicon Valley keeps rediscovering: raw scale is not a strategy, and neither is vibe-based prompting. The win is in the recipe, the data, and the test-time machinery all fitting together. Nice to see an actual system beat the hype parade for once.

Read more about this at: Hugging Face

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.