When a Strong Model Has to Give Advice
Standalone model performance is not fully portable across roles: when a model works through another agent, the receiver, task, and communication protocol can change how much value the model adds to the final work.
Claude Opus 4.8 did very well on a natural-gas assignment until it was told not to do the assignment.
The task came from a new benchmark called CentaurBench. A model had to prepare a short client brief on forces likely to move U.S. natural-gas prices in 2026. When Claude wrote the brief itself, it ranked first among the models tested on that task. Its mean rank across repeated runs was 2.05.
Then the researchers put another model between Claude and the client.
Claude could no longer write the brief. It had to send a short set of process instructions to GPT-3.5-Turbo, which would write the final answer. In that arrangement, Claude’s mean rank fell to 8.15.
Nothing about those two numbers proves that Claude is a bad teacher, or that good solvers make bad coaches. The experiment is too narrow for either conclusion. It does reveal something easier to miss: a model can keep all of its underlying capabilities and still lose much of its value when its output has to pass through another decision-maker.
The question is no longer only, How good is this model? It is also, What happens after this model speaks?
The extra step changes the job
Most familiar AI benchmarks ask a model to produce an answer and then score that answer. The path is direct. The model interprets the task, makes its choices, and hands over the finished work.
CentaurBench adds a second path. In its augmentation condition, an upstream model produces guidance; a fixed GPT-3.5-Turbo worker receives that guidance and produces the deliverable. The upstream model never gets to repair the final answer itself.
That makes the quality of the guidance depend on more than what the upstream model knows. The message has to draw attention to the right constraints. It has to be useful at the level of detail the worker can act on. It has to avoid displacing something the worker would otherwise have done correctly. Then the worker has to interpret it well enough to improve the final answer.
The benchmark therefore changes the object being measured. A standalone test measures what a model can produce. The second test measures what changes when another model receives its advice.
Across the nine assistant models that appeared in both conditions, the direct-work and assistance rankings had a Spearman correlation of 0.48. With only nine models, that estimate was not statistically distinguishable from zero. More revealing is how much the relationship varied by task. On tax preparation, the two rankings were closely aligned, with a correlation of 0.85. On travel planning, the correlation was essentially absent. The top-ranked model changed between direct work and assistance on five of seven tasks.
That variation matters because it rules out the easiest story. There is no clean second leaderboard for a hidden ability called “helping.” GPT-5-Mini was strong in both roles. Tax preparation mostly preserved the ordering. The same model can transfer well in one task and poorly in another.
What moved was not merely the model. The arrangement moved around it.
Sometimes the best help was no help
CentaurBench included one especially useful comparison: GPT-3.5-Turbo also completed each task without any upstream guidance.
On operations research, tax preparation, and travel planning, the unaided worker ranked ahead of every assisted version of itself. Averaged across all seven tasks, only guidance from GPT-5-Mini produced a better mean rank than no guidance.
That sounds stranger than it is. Advice is not an extra processor that can simply be bolted onto a system. Advice changes the next actor’s behavior, which means it can improve the work, do nothing, or make the work worse.
The researchers found examples of each mechanism in the messages themselves. Some weak guidance mostly repeated constraints the worker already had. One low-ranked operations-research scaffold went further: it told the worker to avoid specific solutions and task-specific analysis even though the assignment required practical solutions and an analysis of their trade-offs. The assistance note had pointed attention away from part of the job.
This is the important distinction. A model can contain useful knowledge without causing a useful change downstream.
That idea travels beyond this benchmark more easily than the player-and-coach metaphor does. In a planner-worker agent, a beautifully reasoned plan is valuable only if the worker can execute it. In a review system, a sophisticated critique is useful only if it directs the next pass toward the defects that matter. The relevant quantity is not how impressive the intermediate artifact looks. It is the difference it makes to the final work.
But CentaurBench also contains the reason to resist turning that observation into a law.
The benchmark makes helping unusually hard
The assistance protocol is deliberately restrictive. On the natural-gas task, the upstream model gets 200 to 250 words. It must provide a requirements check, an execution plan, and a final checklist. It is explicitly forbidden to give topic-specific content, examples, recommendations, or sentences that could be copied into the answer.
The paper’s authors describe this as a conservative lower bound on augmentation. A model that is excellent because it knows a great deal about energy markets is not allowed to pass that domain knowledge directly to the worker. The experiment isolates process guidance, which makes the comparison cleaner but the meaning narrower.
The receiver never changes either. Every assistant talks to GPT-3.5-Turbo. So the benchmark cannot tell us whether Claude’s weak showing on the natural-gas task is a stable property of Claude-as-helper, a bad fit with GPT-3.5, a bad fit with the process-only instruction, or some combination of the three.
There is another source of uncertainty. The final outputs are judged by language models rather than by domain experts. Across comparisons seen by at least two eligible judges, the judges picked the same winner 71 percent of the time. In the augmentation condition, agreement was lower: 67.8 percent. The ranking shifts are real results of the published evaluation procedure, but their magnitude still needs external validation.
The strongest rival explanation is therefore not that role does not matter. It is that this particular protocol makes role matter more than it would in a richer, more realistic collaboration.
Other research shows why that rival deserves respect. In a 2025 study of repository-level code generation, researchers tested ways to combine stronger and weaker models on software tasks. Their best pipeline matched the tested strong model’s performance at about 60 percent of its cost. A strong upstream model was not a liability there; the architecture made its contribution useful.
That result kills the catchier thesis. The best worker is not generally the worst helper.
It also sharpens the better one: helpfulness is produced by a relationship between components, not supplied by one component in isolation.
Change the receiver, and the test should change
If that explanation is right, it makes a prediction. Change the receiver or the communication rule, and some assistant rankings should move again.
CentaurBench has not yet run that full test. But two independent lines of evidence make the prediction plausible.
In ChatBench, researchers turned 396 benchmark questions into conversations between people and two AI systems. Across 7,336 user-AI conversations, the model’s accuracy when answering alone did not reliably predict the accuracy of the human-AI pair in several subjects. Once a person entered the causal chain, the original model score no longer described the whole system.
A NeurIPS study of stronger teacher models and weaker student models found something more specific. Strong teachers could improve weaker students, but explanations tailored to a particular student outperformed unpersonalized ones. The same information source became more useful when the message was adapted to the receiver.
Neither study reproduces CentaurBench’s exact experiment. Human conversation is not one-shot model scaffolding, and teaching can aim at learning rather than a single deliverable. That is why they are useful here: they test the broader mechanism under different conditions. In both, performance changes when information has to be interpreted by someone else.
A large 2024 meta-analysis of human-AI experiments makes the boundary equally clear. Across 106 experiments, human-AI combinations performed worse on average than the better solo party, but the results were extremely heterogeneous. Task type and the relative strengths of the human and AI mattered substantially. “Add AI” was not an intervention with one stable effect.
Neither is “add a stronger model.”
A leaderboard for the role, not just the model
This changes a practical evaluation question.
Imagine a team building a system with an expensive planner and a cheaper execution model. A normal model comparison asks which expensive model performs best on a set of tasks. A deployment-shaped comparison asks which planner causes the cheap worker to produce the best final work.
The test is straightforward. Keep the worker fixed. Keep the tasks fixed. Keep the interface fixed. Measure the worker alone. Then place each candidate planner upstream and measure the final output again.
After that, change one thing at a time. Give the planner permission to include worked examples. Let it revise the worker’s answer. Swap in a different worker. Move the same model from planning to review. If the rankings remain stable, standalone strength is traveling well. If they move, the team has learned where the value actually lives.
This does not make ordinary leaderboards useless. A direct score is still evidence about what a model can do directly, and CentaurBench itself shows cases where the ranking carries over. It simply cannot answer a question it never asked.
On the natural-gas task, Claude’s first score belonged to Claude writing the brief. The second score came after the researchers inserted GPT-3.5 and a 200-word note between Claude and the answer.
The model was still Claude. The thing being measured was no longer Claude alone.
Sources
- CentaurBench: Benchmarking LLM Capabilities on Augmenting vs. Automating Real-World Work Tasks, arXiv working paper, submitted August 19, 2026.
- ChatBench: From Static Benchmarks to Human-AI Evaluation, ACL 2025.
- Can Language Models Teach? Teacher Explanations Improve Student Performance via Personalization, NeurIPS 2023.
- An Empirical Study on Strong-Weak Model Collaboration for Repo-level Code Generation, arXiv preprint, 2025.
- When combinations of humans and AI are useful: A systematic review and meta-analysis, Nature Human Behaviour, 2024.