OpenProblemBench Puts AI Up Against 82 Unsolved Math and Physics Problems

Sanket Chaukiyal

October 10, 2026

TL;DR

  • Researchers built OpenProblemBench, a set of 82 unresolved problems pulled straight from the mathematics and theoretical physics literature, not textbooks.
  • Four separate AI evaluator models judge correctness, completeness, and degree of progress since there’s no reference answer to check against.
  • GPT-6-Astra topped the field with a 14.0% mean judged solve rate across seven evaluated configurations.
  • Full-size open models trailed at 5.5-6.7%, and Flash models fell further behind at 2.4-3.7%.

82 Problems No One Has Solved Yet

Every AI benchmark up to this point has shared one quiet assumption: somebody already knows the answer. OpenProblemBench throws that assumption out. The researchers behind it pulled 82 problems directly from the open literature in mathematics and theoretical physics, questions that remain unresolved by the field itself.

That creates an obvious headache. How do you grade a model’s answer when there’s no answer key? The team’s solution was four separate evaluator models, each independently judging a submission on correctness, completeness, and degree of progress, with no reference solution to lean on.

The results are blunt. Across seven evaluated configurations, GPT-6-Astra posted the highest mean judged solve rate at 14.0%. Full-size open models landed at 5.5-6.7%. Flash variants, the smaller and faster models in most lineups, managed only 2.4-3.7%.

Fourteen percent. On problems actual mathematicians and physicists haven’t cracked.

Why a 14% Score Deserves Attention

I’ve sat through a lot of benchmark announcements that amount to a model shaving half a point off a leaderboard nobody outside AI Twitter cares about. This one is different, and I think it deserves more attention than it’s getting: a double-digit solve rate on problems that have stumped professional researchers isn’t a rounding error, it’s a signal.

Think of it like handing a lockpicking kit to someone and asking them to open a vault that the bank’s own engineers haven’t finished designing yet. Most people fumble entirely. A few manage to jiggle the mechanism and get partway in. That’s roughly what a 14% judged solve rate on genuinely open problems looks like: not full solutions walking out the door, but measurable partial progress on locks nobody has fully mapped.

The gap between GPT-6-Astra and the open models is the part worth sitting with. 14.0% versus 5.5-6.7% for full-size open competitors isn’t a photo finish, it’s closer to a two-to-one margin. And Flash models, the leaner versions companies ship when speed and cost matter more than raw capability, fell even further behind at 2.4-3.7%. That tiered gap tells you something about where the frontier actually sits right now: raw scale and training depth still seem to matter enormously once a problem has no known answer to pattern-match against.

Is 14% good? Depends what you compare it to. Against a human mathematician who’s spent years chasing a single conjecture, it’s modest. Against a chatbot, it’s something closer to remarkable. These aren’t trick questions with a clean answer buried in a training set somewhere. They’re the stuff researchers argue about at conferences. A model making any verifiable dent in that is worth more scrutiny than another 95% score on a benchmark everyone’s already memorized.

But fourteen percent also means eighty-six percent of the time, nothing usable comes out. That ratio matters for anyone tempted to treat these systems as research collaborators rather than research assistants.

From Recall Tests to Research Tests

For years, AI benchmarks have mostly tested recall dressed up as reasoning. Math competition problems, coding challenges, trivia with a scientific veneer, all of them share known, checkable answers buried somewhere in the training data or easily derivable from it. Models got very good at these tests, sometimes suspiciously good, and the scores stopped telling researchers much.

OpenProblemBench is part of a broader shift toward evaluation that can’t be gamed by memorization because there’s nothing to memorize. The problems are unresolved. Nobody, human or machine, has the solution yet. That forces evaluators to judge partial progress, logical soundness, and genuine mathematical insight rather than pattern-matching against an answer key.

The four-evaluator-model setup is itself a tacit admission that grading open research is hard even for AI systems built to grade it. Without a ground truth to check against, judgment becomes probabilistic, a consensus of models rating another model’s reasoning. That’s a meaningfully different kind of test than anything resembling a standardized exam, and it’s closer to what actual peer review looks like: messy, contested, and occasionally wrong.

What Comes Next for the Benchmark

Watch whether other labs start reporting their own OpenProblemBench numbers alongside the usual suite of math and coding scores. A 14% solve rate only means something once there’s a wider field of competitors to measure it against, and right now the comparison set is thin.

Keep an eye on how the four evaluator models hold up under scrutiny too. If researchers or outside mathematicians start disputing specific judged solutions, that could either validate the grading approach or expose how shaky AI-judging-AI gets once there’s no answer key underneath it.

And watch the gap between full-size and Flash models. If that spread narrows over the next round of releases, it suggests smaller, cheaper systems are catching up on genuine reasoning rather than just speed. If it widens, it tells a less flattering story about what Flash-tier models are actually built for.

Editor's Note

What gets me about this benchmark isn't the 14% score, it's that anyone bothered to build a test with no answer key at all. I've watched AI benchmarks get memorized into uselessness for years, so a setup where four models have to judge progress on problems nobody's solved feels overdue. I'm skeptical of the judging layer, though. Grading unsolved math with other AI models is a shaky foundation, and I'll want to see independent human mathematicians weigh in before I trust these numbers fully.

– Sanket Chaukiyal, founder, SmartChunks

FAQ

What is OpenProblemBench exactly?

It's a benchmark of 82 unresolved problems pulled from the mathematics and theoretical physics literature, built to test whether AI models can make real progress on questions nobody has solved yet, not just recall known answers.

How do researchers grade answers if the problems are unsolved?

Four separate evaluator models independently judge each submission on correctness, completeness, and degree of progress, since there's no reference solution to check against.

Which model performed best on the benchmark?

GPT-6-Astra led with a mean judged solve rate of 14.0% across seven evaluated configurations, ahead of full-size open models at 5.5-6.7% and Flash models at 2.4-3.7%.

Does a 14% solve rate mean AI is solving open math and physics problems?

Not outright. The score reflects judged partial progress and completeness on unresolved problems, not confirmed full solutions, so it measures incremental reasoning rather than breakthrough proofs.


Source: arXiv

Sanket Chaukiyal — Editor at Smart Chunks

Sanket Chaukiyal

Technology editor • 12+ years in editorial

Sanket is the founder and editor of Smart Chunks. He spent over six years at Autocar India (Haymarket SAC Publishing) as Sub Editor and Senior Copy Editor, and later served as Account Director (Content) at Rite Knowledge Labs. He holds a Master's in Media and Communication from the Symbiosis Institute of Media and Communication.

All articles → • LinkedIn