When the Customer Changes the Test
An interactive benchmark can correctly record task completion without establishing that its simulated user followed the prescribed conditions; checking those conditions and checking their resemblance to real people are separate jobs.
In a telecom-support simulation, a language model playing the customer offered a name and phone number before being asked. The episode, excerpted in UserProxyBench, counted as an agent success. It sounds like a sensible start to a support conversation. There was just one problem: the customer had been instructed to release information only when requested.
The introduction removed an opportunity to observe whether the agent would ask for it. Nothing about identifying yourself to a support agent is inherently wrong. But if an experiment is meant to test whether the agent can discover what information it needs, supplying that information changes the experiment. The customer had done part of the work before the worker could be tested on it. The distinction matters even if the resulting service is excellent. A customer can make an encounter easier without telling us whether the agent has become more capable at handling it.
Ashish Jain and Armaan Sandhu made that distinction the subject of their September 29 paper. Keeping one agent fixed across 375 tasks and seven user-model configurations, they audited the customers separately. Their model-judged audit found user-rule violations in 24.4 percent of scored episodes in which the agent succeeded. A passing customer had to satisfy every applicable check. The same model generated the criteria and judged compliance; human validation was limited.
The task could be completed even though the test had not been delivered as prescribed. That leaves a question the completion score cannot answer: did the agent demonstrate the ability the researchers meant to test, or did the conversation give it a different job?
Who gets to be helpful?
There is a reasonable objection. Real customers do not follow researchers’ scripts. Someone who knows what a support agent needs may volunteer it immediately. Why penalize a simulator for making a conversation go smoothly?
The creators of the original τ-bench, the benchmark family underlying the new work, anticipated this problem. They acknowledged limits in simulated users’ reasoning, memory, and instruction-following. They also offered a defense: actual people vary in knowledge and skill, and agents must handle that variation. An irregular customer is not automatically a defective test.
Their outcome-centered design has a practical virtue. Checking the final database state can establish whether a requested change happened without requiring one approved conversation. An agent should be able to find a better route. Otherwise an evaluation risks rewarding obedience to a script over successful assistance.
Freedom to find a route, though, is different from freedom to change the starting conditions. A test can permit many good solutions while specifying what the customer knows and when they reveal it. The agent can choose its questions; the simulator can still be required to wait for them. Enforcing that requirement tells us whether the experiment followed its own rules. It cannot tell us whether those rules describe real customers.
Preethi Seshadri and colleagues approached that question in Lost in Simulation. With a GPT-4o agent held fixed, changing the models playing its retail customers produced mean success rates from 67.0 to 75.9 percent across 115 tasks and three runs. Their separate comparison with US participants complicated any simple story about artificially helpful customers: simulation underestimated success on the hardest tasks and overestimated it on moderately difficult ones. The human comparison used a selected, difficulty-balanced subset, with participants role-playing assigned customers.
A simulator can make an agent look better or worse. Choosing the less flattering score does not necessarily bring an evaluation closer to people. The researchers’ participant instructions supply an especially revealing detail: people were explicitly told to begin by supplying their assigned email address. Early identification was part of that experiment’s design.
The two introductions belong to different tasks. Comparing them cannot establish which disclosure policy is more realistic. It does show why “the user volunteered information” is an incomplete diagnosis. In the human study, withholding identity would frustrate the setup; in the telecom simulation, offering it early frustrated the setup. We need to know the instruction before we can judge the behavior.
The instructions need checking, too
Adding another score might seem to settle this: one number for completion, another for the customer’s obedience. But the second score is also a measurement. It can miss something that matters.
The new paper supplies a useful check on itself. A separate Telecom inspection of user-side tool actions found 54 of 658 examined conversations missing a required action. In 23 of them, every applicable language-model check had passed. A favorable audit could coexist with a missing operation. Passing the language-model audit had not reliably established whether the required action occurred.
For someone comparing agents, the practical consequence is that the conversation belongs with the result. A completion score describes achievement under the conditions actually supplied. Before treating a shorter interaction as evidence of less work, we need to know whether information-gathering was performed more efficiently or made unnecessary. Before calling those conditions realistic, we need a comparison with people.
The helpful introduction is still perfectly plausible customer service. But before crediting an agent for knowing what to ask, look at what the customer had already said.
Sources
- Ashish Jain and Armaan Sandhu, UserProxyBench: Evaluating LLM User Simulators for Agent Benchmarks and Training, September 29, 2026, version 1
- Preethi Seshadri and colleagues, Lost in Simulation: LLM-Simulated Users are Unreliable Proxies for Human Users in Agentic Evaluations, January 23, 2026, version 1
- Shunyu Yao and colleagues, τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains, June 2024, version 1