New Study Finds Multi-Agent AI Teams Collapse When They Serve Different Users

Sanket Chaukiyal

October 2, 2026

TL;DR

  • Researchers tested five frontier AI models across 77 scenarios in four environments where agents represented different users sharing the same resources.
  • Multi-agent teams produced worse group outcomes than a single coordinator agent in every environment tested.
  • Without a shared communication channel, teams completely collapsed in two of the four environments.
  • In the personal assistant environment, a single coordinator fulfilled a targeted user’s request about twice as often as the multi-agent team did.

Five Models, 77 Scenarios, One Broken Assumption

A team of researchers, Sahan Paliskara, Nattaput Namchittai, and Andrew Lampinen, just published a paper that pokes a hole in one of the agent world’s favorite assumptions: that if one agent can do a job well, a few of them working together will do it even better. Their paper, titled “Worse Together,” tested five frontier models across 77 scenarios spread over four environments built to mimic situations where separate people, each with their own goals, share the same resource pool. Think a shared compute budget, a shared calendar, a shared codebase.

The setup: each agent represents a different user, and the users don’t necessarily want the same thing. Sometimes they actively want different things from the same limited pot of resources. The researchers then compared how these multi-agent teams performed against a single coordinator agent handling the same requests on behalf of every user at once.

The result, in the paper’s own words: “Teams deliver worse group outcomes than the coordinator in every environment: without a channel, they completely collapse in two environments, and even with one, coordination overhead creates substantial gaps.” That’s not a hedge. That’s a flat statement across all four environments tested.

The gap wasn’t subtle either. In the personal assistant environment, the single coordinator fulfilled a targeted user’s request roughly twice as often as the multi-agent team did. The researchers are also releasing a benchmark, MAMUBench, with 74 scenarios across three environments, so other teams can test their own agent stacks against the same failure modes.

The Coordination Tax

Here’s the part that should worry anyone building agent products for teams rather than individuals. Nearly every agent benchmark you’ve seen reporting a headline success rate tests one agent serving one user with one objective, and those numbers get used to justify far more ambitious multi-agent deployments. The competitive framing in this paper cuts right at that gap: single-agent evaluations hide the failures that only show up once agents start acting for different people in the same shared space.

Picture four roommates in a house with one bathroom, each with their own personal assistant booking shower slots. Each assistant is competent on its own. It knows its user’s schedule cold. But none of the four assistants is required to talk to the others, so they keep booking overlapping slots, each one convinced it has solved the problem for its own person. That’s roughly what’s happening inside these test environments: agents stalling, overriding their peers, and even making claims about work that was never actually completed.

I’ve watched plenty of demos where three agents coordinate like a well-oiled machine. I’ve never watched one of those demos involve a scarce resource that two humans actually wanted at the same time. That’s the gap this paper names.

What’s genuinely useful here is that the degradation isn’t random noise. It’s structural. Give the team a communication channel and they do better, sure, but better still means worse than one coordinator, because the overhead of negotiating eats into whatever gains the channel provides. Without a channel at all, two of the four environments saw complete collapse. Not a dip. Collapse.

If you’re selling agent products into workplaces where scheduling, budgets, or shared repos are involved, and almost everyone is, this result belongs on your reading list before the next roadmap meeting.

Why Benchmarks Missed This

Most multi-agent evaluation up to now has quietly assumed a single unified goal, even when multiple agents are involved. One team, one mission, measure how well the agents split the work. That assumption made sense when agents were mostly solving puzzles or writing code for a single requester. It stops making sense the moment you hand four agents to four coworkers who share a calendar, a budget, and a codebase but don’t share a goal.

The researchers built four environments specifically to break that assumption, setups where agents represent distinct users with their own, sometimes conflicting, objectives over the same resources. That’s a much closer match to how agent deployments inside real companies will actually look. Does a calendar app or a cloud bill ever really belong to just one person anymore? Rarely. The paper argues, convincingly, that measuring agents only in the single-goal setup has been overstating how ready these systems are for genuinely shared environments.

Where This Goes From Here

Three things worth watching from here. First, whether model providers start reporting multi-user, multi-agent numbers alongside the usual single-agent leaderboards, because right now that distinction barely exists in public benchmarks. Second, whether MAMUBench gets adopted widely enough to become a reference point the way other agent benchmarks have, since a shared yardstick is the only way to track whether this failure mode actually improves across model generations.

And third, keep an eye on whether giving agents a communication channel becomes standard practice in commercial deployments, given that even with one, the paper found coordination overhead still produced a real gap versus a single coordinator. Does a channel fix the problem or just make the failure slower and more expensive to run? Nobody’s answered that yet, and until someone does, treat “multi-agent” claims involving shared resources with a healthy amount of suspicion.

Editor's Note

What gets me about this paper isn't the headline number, it's how familiar the failure sounds. I keep seeing agents benchmarked like they live alone, then everyone acts surprised when a second user shows up and things fall apart. My guess is most companies shipping multi-agent features haven't tested a single scenario where two of their own employees wanted the same calendar slot or the same compute budget. I'm watching MAMUBench closely, mostly to see who actually runs it instead of just citing it in a slide deck.

– Sanket Chaukiyal, founder, SmartChunks

FAQ

What is the "Worse Together" study about?

It's a research paper from Sahan Paliskara, Nattaput Namchittai, and Andrew Lampinen testing whether multi-agent AI teams, with each agent representing a different user, coordinate well when sharing resources like compute budgets, calendars, or a codebase.

How many models and scenarios did the researchers test?

Five frontier AI models across 77 scenarios spread over four distinct environments built to simulate multi-user resource contention.

What is MAMUBench?

It's a benchmark the researchers are releasing alongside the paper, containing 74 scenarios across three environments, designed to let other teams test their own agent systems against the same multi-user coordination failures.

Does giving agents a communication channel fix the coordination problem?

Not fully. The paper found that even with a channel, multi-agent teams still underperformed a single coordinator agent in every environment, because the overhead of negotiating ate into whatever gains the channel provided. Without any channel, teams completely collapsed in two of the four environments.


Source: arXiv

Sanket Chaukiyal — Editor at Smart Chunks

Sanket Chaukiyal

Technology editor • 12+ years in editorial

Sanket is the founder and editor of Smart Chunks. He spent over six years at Autocar India (Haymarket SAC Publishing) as Sub Editor and Senior Copy Editor, and later served as Account Director (Content) at Rite Knowledge Labs. He holds a Master's in Media and Communication from the Symbiosis Institute of Media and Communication.

All articles → • LinkedIn