GAUGE Study Finds LLM Judges Can’t Tell Which AI Agents Actually Finish the Job

Sanket Chaukiyal

September 14, 2026

TL;DR

  • 57.5% of AI agent conversations rated “satisfied” by a blind panel actually failed the customer’s task.
  • Judge disagreement on close agent matchups hit 31%, up from under 1% on lopsided pairs.
  • The study, called GAUGE, ran 25 task-oriented agents from 6 providers through tau^2-bench and SimulatorArena.
  • Researchers Umesh Bodhwani, Thanh Tran, and Kai Wei built the protocol specifically to audit this evaluation pattern, not to promote a new agent.

Inside the GAUGE Protocol

Three researchers, Umesh Bodhwani, Thanh Tran, and Kai Wei, built an offline evaluation protocol called GAUGE to check whether a now-standard grading method for AI agents actually measures what it claims to. That method usually works like this: a simulated persona plays the customer, a candidate agent handles the conversation, and a separate LLM judge reads the transcript afterward and decides whether the customer would have walked away happy.

GAUGE ran that pipeline against 25 task-oriented agents pulled from six different providers, tested across two established benchmarks, tau^2-bench and SimulatorArena. Then the researchers checked the judge’s verdicts against ground truth: did the agent actually complete the task the simulated customer asked for, verified independently of the satisfaction score.

The gap was not small. In 57.5% of conversations that a blind panel rated as satisfying, the agent had failed the customer’s actual task. The researchers didn’t hedge on this. Their own phrasing: satisfaction “carries essentially no information about task success, as conversations rated satisfied by our blind panel are decorrelated from actual success, with 57.5% of them failing the customer’s task.”

The judging pipeline also got noticeably shakier once the agents being compared were close in ability. Disagreement on which agent performed better sat under 1% when one agent was clearly stronger. On tightly matched pairs, that disagreement jumped to 31%.

What This Means

Here’s the part that should bother anyone using an LLM judge as a launch gate: the failure isn’t random noise, it’s structural. When one agent is obviously better, the judge nails it almost every time. That’s not the hard case though. The hard case is choosing between two agents that are both decent, and that’s exactly where the pipeline falls apart, disagreement jumping to 31%.

Picture a bathroom scale that reads perfectly accurate when you’re comparing a bowling ball to a feather, but starts wobbling by five pounds the second you put two similar-sized rocks on it. The tool works fine for the comparisons nobody needed help with. It breaks exactly where the decision actually mattered.

I’ve sat through enough agent review meetings to know how tempting it is to treat a high satisfaction score as a green light for shipping. This paper is a fairly blunt reminder that “the customer seemed happy” and “the customer got what they asked for” are not the same finding, and treating them as interchangeable is how a company ends up promoting an agent that talks a good game while quietly failing its users.

Why does this happen at all? Because the simulated judge is grading tone, helpfulness, and whether the chat felt resolved, not whether anything actually got done. An agent can apologize gracefully, sound competent, and wrap the conversation on a warm note without ever having booked the flight, issued the refund, or cancelled the subscription it was asked to handle. If a satisfaction score can’t tell you whether the task got done, what exactly is it measuring?

How We Got Stuck Grading Vibes

Task-oriented agents, the kind that handle customer service, bookings, refunds, and account changes, aren’t reviewed by humans one conversation at a time anymore. There are too many candidate models, too many prompt tweaks, too many fine-tunes to check by hand. So the industry built a shortcut: a simulated persona plays the customer, the candidate agent responds, and a separate LLM scores the transcript.

This pattern, often called LLM-as-a-judge, spread fast because it’s cheap and it scales. tau^2-bench and SimulatorArena, the two benchmarks GAUGE tested against, belong to that same wave of synthetic evaluation infrastructure that companies now lean on as an offline gate before an agent ever reaches a real customer.

GAUGE doesn’t argue that simulated judging is worthless. It argues that nobody had rigorously checked whether the judge’s verdict tracks the thing it’s supposed to measure, real task completion verified against ground truth rather than vibes. That check hadn’t really existed at this scale, across 25 agents and six providers, until this paper ran it.

Signals That Will Tell Us If This Sticks

Watch whether companies running these evaluation gates start pairing satisfaction scores with a hard, ground-truth task-completion check instead of leaning on the judge’s verdict alone. Watch too for other teams trying to reproduce that 31% disagreement figure on their own agent lineups. If it holds up outside tau^2-bench and SimulatorArena, a lot of published agent leaderboards are overdue for a second look.

And keep an eye on whether the GAUGE authors, or anyone else, follow this diagnosis with a fix. Flagging a broken gate is only half the job. The harder question is what replaces the satisfaction score when the whole point of the simulator was to avoid paying humans to check every conversation by hand.

Editor's Note

What gets me about this paper is how quiet the failure mode is. No agent looks obviously broken, the judge just can't tell two good ones apart, and that's arguably worse than an outright crash. I've watched teams treat a satisfaction dashboard as proof of readiness, and 57.5% of "satisfied" chats actually failing the task should end that habit fast. My guess is that within a year, serious evaluation pipelines pair LLM judges with a hard completion check by default, not as an afterthought bolted on later.

– Sanket Chaukiyal, founder, SmartChunks

FAQ

What is GAUGE?

GAUGE is an offline evaluation protocol built by researchers Umesh Bodhwani, Thanh Tran, and Kai Wei to test whether user-simulator and LLM-as-a-judge pipelines actually rank task-oriented AI agents correctly. It compares satisfaction ratings from a simulated customer against verified, ground-truth task completion.

What did the study find about satisfaction scores?

The researchers found that 57.5% of conversations a blind panel rated as satisfying actually ended in the agent failing the customer's task. In their words, satisfaction "carries essentially no information about task success."

How many agents and providers were tested?

The study covered 25 task-oriented agents from six different providers, evaluated across two existing benchmarks, tau^2-bench and SimulatorArena.

Why does the judge disagreement rate matter?

When comparing agents with a clear performance gap, the judging pipeline disagreed on the outcome less than 1% of the time. But comparing two closely matched, strong agents, disagreement jumped to 31%, meaning the tool is least reliable exactly when a real decision is on the line.


Source: arXiv

Sanket Chaukiyal — Editor at Smart Chunks

Sanket Chaukiyal

Technology editor • 12+ years in editorial

Sanket is the founder and editor of Smart Chunks. He spent over six years at Autocar India (Haymarket SAC Publishing) as Sub Editor and Senior Copy Editor, and later served as Account Director (Content) at Rite Knowledge Labs. He holds a Master's in Media and Communication from the Symbiosis Institute of Media and Communication.

All articles → LinkedIn