Inside ‘The Station’: AI Agents Found New Math With No One in Charge

Sanket Chaukiyal

August 26, 2026

TL;DR

  • Researchers built “the Station,” an open-world environment where AI agents from different model families work on math problems with no central boss and no scripted pipeline.
  • Across 14 problems, 12 pulled from the AlphaEvolve catalogue plus two case studies, the agents landed novel results on five of them.
  • The headline find: new exact 604-point kissing configurations in dimension 11, plus an improved lower bound for Erdős’s minimum-overlap problem.
  • Unlike earlier automated discovery tools that hand back raw numbers, these agents wrote proofs and explanations a working mathematician can actually check.

Agents, No Boss, No Script

No coordinator. No script. That’s the entire premise of a paper from Stephen Chung, Wenyu Du, and William J. Wesley, and it’s almost stubbornly simple as ideas go: what happens if nobody tells the AI what to work on next?

Most automated discovery systems run on a script. Someone defines the search space, sets an objective, and the model grinds until it hits a number worth reporting. The Station skips that step entirely. The researchers describe it as “an open-world multi-agent environment in which AI agents from different model families pursue a shared research goal without a central coordinator or scripted pipeline.” Agents pick their own direction, run their own experiments, and read each other’s work as it piles up, building something closer to a shared lab notebook than a single model’s output log.

The test set covered 14 problems: 12 construction problems lifted from the AlphaEvolve catalogue, a known benchmark, plus two additional case studies the researchers added themselves. On five of those fourteen, the agents produced results novel relative to the existing literature, not just faster reproductions of what was already known.

The standout is a set of exact 604-point kissing configurations in dimension 11, a geometry problem about how many spheres can touch a central sphere without overlapping. The agents also improved the lower bound on Erdős’s minimum-overlap problem, a decades-old question in additive combinatorics, and turned up new finite-field Kakeya sets along the way.

Our Take

Here’s what actually matters in this paper, and it isn’t the kissing number. It’s that the agents didn’t just spit out a configuration and call it done. According to the researchers, “the agents produced not only numerical constructions but also theorems and analyses explaining how those constructions work, making the results more interpretable and easier for mathematicians to build upon.” That’s the gap between a black box handing you an answer and a colleague handing you a proof you can actually argue with.

I’ve read enough “AI discovers new math” headlines to be reflexively skeptical of them, and most turn out to be a search algorithm finding a marginally better number inside a space a human already boxed in. This paper reads differently, mainly because the coordination problem it’s solving is genuinely hard. Getting multiple models from different families to divide labor, avoid duplicated effort, and build on each other’s partial results without a project manager looks a lot less like running an experiment and a lot more like running an open-source community.

Picture a construction site with five crews from five different companies, no blueprint, and no site manager standing there with a clipboard. Normally you’d expect chaos, or at best five disconnected sheds. Instead, according to the paper, five load-bearing walls went up, and each crew left behind a set of notes explaining exactly why their wall doesn’t fall down.

And that’s before you weigh the comparison the researchers are clearly inviting. AlphaEvolve, the system whose problem catalogue supplied 12 of the 14 test cases here, is a single coordinated search process built around one architecture. The Station takes AlphaEvolve’s own benchmark and clears a chunk of it using a fundamentally different setup: multiple models, no shared owner, no scripted objective function. Choosing those exact 12 problems as the test bed is not a neutral move, and it’s a pointed one.

Does decentralization actually scale past a curated 14-problem set, or does it just work well on the problems someone was confident enough to publish? That’s the real open question, and the paper doesn’t pretend otherwise.

Where This Sits

Automated mathematical discovery has mostly lived inside tightly scripted pipelines: define the search space, pick an objective, let one model or one evolutionary process iterate until something scores well. AlphaEvolve is the most visible recent example of that approach, and it’s exactly the benchmark this paper borrows from.

What’s different about the Station isn’t the math, it’s the org chart. There’s no central coordinator assigning problems or approving results before they count. Agents from different model families choose their own directions, which is a small thing to write down and a genuinely hard thing to get working without everyone duplicating effort or wandering off into dead ends.

But the interpretability piece is arguably the bigger deal long term. A raw numerical construction is useful for exactly one thing: the number. A construction plus a proof explaining why it works is something a mathematician can extend, cite, and build a follow-up paper on. That’s the difference between a calculator and a collaborator, and it’s why this paper is getting attention beyond the specific kissing number in dimension 11.

What to Watch Next

The obvious next step is peer review and independent verification. A 604-point kissing configuration in dimension 11 is a specific, checkable claim, and the mathematical community will either confirm it holds against known bounds or find a flaw. Watch for whether other researchers replicate the construction or the improved Erdős minimum-overlap bound.

Also worth tracking: whether this approach scales beyond a curated 14-problem set the authors picked themselves, which is reasonable for a first paper but leaves open how the Station performs on problems nobody chose in advance. Is a hand-picked benchmark enough to prove decentralized collaboration works, or just enough to make a compelling paper?

Keep an eye, too, on whether other labs pick up the coordinator-free, multi-model approach or whether this stays a one-off proof of concept. If a second team reproduces even one of these five results with a different mix of models, that’s the signal this idea has legs beyond its own paper.

Editor's Note

What gets me about this paper isn't the kissing number, it's that nobody was in charge. I've watched multi-agent demos collapse into models stepping on each other's work more times than I can count, and here five separate families apparently didn't. I'm holding off on full enthusiasm until independent mathematicians check that dimension 11 construction, but if the coordination trick holds up on harder, unpicked problems, that's the part I'd bet on.

– Sanket Chaukiyal, founder, SmartChunks

FAQ

What is "the Station"?

It's an open-world multi-agent environment where AI agents from different model families work on math problems together without a central coordinator or a scripted pipeline, described in a paper by Stephen Chung, Wenyu Du, and William J. Wesley.

How many problems did the agents actually solve?

The team tested 14 problems total, 12 drawn from the AlphaEvolve catalogue and two additional case studies. The agents produced results novel relative to prior literature on five of those fourteen.

What is the 604-point kissing configuration about?

It's a geometry result in dimension 11 describing how spheres can be arranged to touch a central sphere without overlapping. The Station's agents found an exact configuration with 604 points, one of the paper's headline results.

How is this different from AlphaEvolve?

AlphaEvolve is a coordinated search system built around a single architecture. The Station borrowed 12 of AlphaEvolve's own benchmark problems but tackled them with multiple AI models operating independently, without a central coordinator, and produced human-readable proofs alongside the raw constructions.


Source: arXiv

Sanket Chaukiyal — Editor at Smart Chunks

Sanket Chaukiyal

Technology editor • 12+ years in editorial

Sanket is the founder and editor of Smart Chunks. He spent over six years at Autocar India (Haymarket SAC Publishing) as Sub Editor and Senior Copy Editor, and later served as Account Director (Content) at Rite Knowledge Labs. He holds a Master's in Media and Communication from the Symbiosis Institute of Media and Communication.

All articles → LinkedIn