TL;DR
- NVIDIA fine-tuned its Nemotron 3 model family to hit gold-medal thresholds on both the 2026 International Olympiad in Informatics (IOI) and the 2026 International Mathematical Olympiad (IMO).
- Nemotron-3-Ultra-CC scored 535.4 out of 600 on IOI 2026, clearing the gold cutoff of 361.12 and beating the top human contestant’s 498.27.
- A separate Nemotron 3 Ultra system scored 30 out of 42 on IMO 2026, passing the official gold threshold of 29 using a generate-verify-refine pipeline.
- The IOI result came from an unofficial, prospective run, not part of the official human standings, and NVIDIA has released the model, checkpoints, datasets, and inference code publicly.
The Numbers Behind Two Gold Medals
NVIDIA published a report on Hugging Face on October 7, 2026, laying out how its Nemotron 3 family was fine-tuned to reach gold-level performance in two competitions that rarely get mentioned in the same sentence. One tests algorithm design and code execution under brutal time constraints. The other demands formal, natural-language mathematical proofs. Nemotron did both.
On the IOI 2026 benchmark, Nemotron-3-Ultra-CC, a model with 550 billion total parameters and 55 billion active parameters, scored 535.4 out of 600 using a technique the team calls GenCorrect. That’s not just above the gold threshold of 361.12. It’s higher than the best score any human contestant posted, 498.27.
The IMO side used a different setup: Nemotron 3 Ultra specialists running inside a generate-verify-refine pipeline, where the model drafts a proof, checks it, and revises before submitting. That system landed 30 out of 42 points, just past the official gold cutoff of 29. Getting there took a purpose-built training corpus of 414,890 quality-filtered examples drawn from 15,818 proof problems, fed through supervised fine-tuning and reinforcement learning.
NVIDIA’s own researchers are blunt about where the credit goes. As they put it: “The medals were not produced by fine-tuning alone, and they were not produced by brute-force sampling alone. They came from co-designing the model, the data, and the inference loop.” Worth noting: the IOI run was unofficial and prospective, meaning it never sat alongside the actual human competitors’ submissions in a sanctioned judging process. NVIDIA has released the Nemotron-3-Ultra-CC model, the IMO 2026 SFT and RL checkpoints, the training datasets, and the inference pipelines through Hugging Face and GitHub, so outside researchers can poke at the claims themselves.
The Real Trick Isn’t the Model, It’s the Loop
I’ve read a lot of benchmark press releases this year, and most of them read like marketing copy wearing a lab coat. This one at least ships its receipts: exact scores, exact thresholds, exact dataset sizes, and a model you can actually download and test. That alone puts it ahead of a lot of “state of the art” claims that evaporate the moment someone asks for the weights.
Here’s the thing that actually matters, though. Nobody handed a stock Nemotron model a coding problem and a proof problem and watched it ace both cold. NVIDIA built separate training recipes, separate data pipelines, and separate inference strategies for each domain, then bolted them onto the same base architecture. Picture a long-distance runner who also wins judo tournaments, not because their body magically adapted overnight, but because a coaching staff rebuilt the diet, the drills, and the warm-up ritual from scratch for each sport, using the same athlete as the raw material. That’s closer to what happened here than any story about a single model “becoming smarter.”
So does beating the top human IOI score in an unofficial, prospective run actually prove the model would medal in a real, judged competition? Not quite, and NVIDIA doesn’t claim it does. The result sits outside the official standings for a reason. And if a model needs a bespoke 414,890-example training corpus and a custom generate-verify-refine pipeline just to clear one domain’s gold bar, how much does that tell us about general reasoning versus very well-engineered specialization? That question doesn’t have a clean answer yet, and NVIDIA’s own framing, crediting the model, the data, and the inference loop together, suggests even they see it as an engineering achievement as much as an intelligence one.
There’s also a quieter signal buried in the release strategy. NVIDIA didn’t just announce numbers and walk away. It published the model weights, both sets of checkpoints, the datasets, and the inference code. Labs that chase Olympiad-style headlines without opening up their pipelines for inspection are making a very different bet than NVIDIA just made.
Why IOI and IMO Don’t Usually Mix
The IOI and IMO have historically rewarded very different kinds of thinking. IOI contestants write and execute real code against strict time and memory limits, where a program either passes the test cases or it doesn’t. IMO contestants write formal proofs in natural mathematical language, where partial credit, rigor, and elegant reasoning all count toward a score.
Training a model that’s genuinely strong at both has been treated as a harder problem than making a model good at either one alone, since the skills reward different failure modes: a coding model that times out fails completely, while a proof-writing model that stumbles on one lemma can still pick up partial points. NVIDIA’s own background framing leans into this gap explicitly, describing IOI as testing algorithm design and code execution under strict constraints, and IMO as requiring formal natural-language proofs.
That’s what makes the dual result noteworthy rather than just another leaderboard entry. But it also explains why the IOI score carries an asterisk. Unofficial, prospective evaluations let a company run its model against problems from a real competition without the scrutiny, adversarial judging, and procedural rigor that comes with an official standing. The IMO result, scored against the official gold threshold of 29, reads as somewhat more comparable to how human medalists are actually judged.
What Comes Next
Watch for whether IOI organizers or independent researchers ever formally validate a comparable run under official competition conditions, since an unofficial score beating the top human contestant is a very different claim than a sanctioned one. Keep an eye on whether outside labs, given the released Nemotron-3-Ultra-CC weights and IMO checkpoints, can actually reproduce the 535.4 and 30-point scores, because open weights invite exactly that kind of scrutiny. And watch whether other model families attempt the same dual-domain gold push in 2027, since NVIDIA just published a rough blueprint, SFT plus RL plus a generate-verify-refine loop, for anyone willing to build the specialized data pipelines to match it.
Editor's Note
What gets me about this one isn't the scores, it's that Nvidia actually shipped the receipts: weights, checkpoints, datasets, all of it. I'd rather see one honest unofficial result with open code than three polished claims nobody can check. I'm watching whether the IOI number holds up once someone outside Nvidia runs it under real judging conditions. My guess is it mostly does, but not by the same margin.
– Sanket Chaukiyal, founder, SmartChunks
FAQ
What exactly did NVIDIA claim about Nemotron 3?
That fine-tuned versions of the Nemotron 3 family reached gold-medal level performance on both the 2026 IOI and the 2026 IMO, using domain-specific post-training and inference pipelines rather than a single generic approach.
How did the IOI 2026 score compare to human competitors?
Nemotron-3-Ultra-CC scored 535.4 out of 600, which is higher than the best human contestant's score of 498.27 and well above the gold threshold of 361.12.
Is the IOI result an official competition finding?
No. It was an unofficial, prospective evaluation run and was not included in the official human IOI standings for 2026.
What did NVIDIA make available to the public?
The Nemotron-3-Ultra-CC model, the IMO 2026 SFT and RL checkpoints, the training datasets used for fine-tuning, and the inference pipelines, all released through Hugging Face and GitHub.
Source: Hugging Face
