Skip to content

When a Second AI Opinion Repeats the First

A second AI reviewer is only partly independent if it sees the first agent's conclusion before forming its own; preserving the complete record and deciding when to reveal prior verdicts are separate choices whose value depends on source reliability and evidence quality.

A second opinion sounds independent because it comes second.

Paul-Peter Arslan’s new experiment shows how easily that intuition can fail.

In a synthetic six-link supply chain, one agent after another handled a stage of the work and filed a structured report. An outside auditor then had to identify where a defect had entered. The reports included observations, but they could also include a compact conclusion: Suspected origin: step N.

Arslan ran a follow-up in which that one clause disappeared. The observations stayed. The report format stayed. The auditor stayed. The task stayed.

When the upstream agents had pointed to the wrong source, removing the conclusion changed the answer dramatically. In episodes built from GPT-4.1 Mini reports, the auditor’s attribution accuracy rose from 4.1 percent to 45.2 percent. In the corresponding Claude Haiku 4.5 condition, it rose from 8.7 percent to 35.4 percent. The paper reports that the first auditor had been overwhelmingly staying inside the set of culprits proposed upstream; once the culprit field disappeared, that adherence collapsed.

The second reviewer had not lost evidence about what happened.

It had lost somebody else’s answer.

A reviewer can inherit a hypothesis

The cleanest way to understand this result is not that provenance is bad. It is that a provenance record can contain different kinds of things.

Some entries are observations: a test failed, a quantity changed, a file was modified, an agent called a tool, a service returned an error.

Other entries are interpretations: this agent caused the failure; that edit is the root cause; this dependency is responsible.

Those can sit next to each other in the same report, but they do not play the same role in a review. An observation gives the reviewer material from which to form a hypothesis. A prior verdict gives the reviewer a hypothesis before it has finished interpreting the material.

That difference matters even when the reviewer is a different model.

Arslan’s study used auditors distinct from the chain models, specifically to reduce same-model coupling. Yet model identity did not make the second judgment independent of the first judgment’s content. The prior diagnosis arrived through the interface.

A separate line of research makes the broader phenomenon hard to dismiss as a quirk of this one synthetic pipeline. Ante Kapetanovic and colleagues gave language-model judges the same items while varying metadata that included a prior score. Across 192,000 attempted evaluations, seven of eight tested models shifted toward the supplied score. In a categorical task with human ground truth, the anchored metadata blocked many corrections and sometimes pulled initially correct judgments toward an assigned wrong label.

The tasks are different. One study is fault attribution; the other is evaluation. But they share a useful mechanism: a previous judgment becomes part of the next judgment’s input.

Calling the second model “independent” does not make that causal link disappear.

More information and more independence are different goals

There is an easy overreaction: if prior conclusions contaminate review, hide them.

The same experiment shows why that rule is too simple.

When the upstream suggestion was correct, deleting the conclusion made the primary auditor worse. In the GPT-chain condition, accuracy fell from 70.5 percent to 55.7 percent. In the Claude-chain condition, it fell from 80.6 percent to 71.6 percent.

A correct diagnosis is useful information.

That turns the problem from a bias story into an interface problem. The conclusion field can act as both an anchor and a compressed signal. Whether showing it helps depends partly on how reliable the source is and how much independent evidence the reviewer already has.

The paper’s broader report-versus-document comparison reveals another complication. In the subset where no upstream agent had named the true source, the auditor reached 60.3 percent accuracy from the raw documentation in one chain condition, far above its 4.1 percent result from the filed reports. Removing the culprit clause raised report-based accuracy to 45.2 percent, but still did not reach the raw-document result.

So the conclusion field was not the whole problem.

The remaining gap is consistent with evidence being lost or compressed in the filed reports, but the study did not validate exactly which missing details caused it.

That possibility has independent support. The ACL 2026 TraceElephant benchmark studies failure attribution under fuller execution observability. Its authors report gains of up to 76.5 percent over a partial-observation counterpart. Their question is not whether a previous diagnosis anchors a reviewer; it is whether the reviewer has enough of the execution record to diagnose the failure in the first place.

Put the two findings together and the review problem has at least two axes:

  • Does the reviewer have enough evidence?
  • Does the reviewer get a chance to interpret that evidence before being shown somebody else’s conclusion?

A system can be weak on either one.

A thin report may leave the auditor ignorant. A rich report that leads with a confident diagnosis may leave it informed but dependent.

The most interesting result was not the preregistered one

This is also where the new study needs to be read carefully.

The paper began with a different confirmatory hypothesis: that collective-responsibility framing would increasingly degrade escalation as the organizational chain grew. That preregistered result was null for both chain models.

The result about the culprit field came later. The paper labels the single-clause deletion exploratory. The originally registered restricted-audit interface also had to be changed because the described inter-agent messages did not exist in the implemented architecture. The author reports that deviation explicitly and places the causal weight on the later matched deletion experiment instead.

That is good research hygiene, but it changes what the evidence can support.

The clause-deletion effect is unusually clean inside this study. It does not establish that real software reviewers will move by forty percentage points when an upstream diagnosis is hidden. Both direct domains are synthetic. The paper is a fresh, single-author preprint. The second domain and additional auditors reproduce the harm from wrong suggestions, but the cost of hiding correct suggestions is not uniform across those follow-ups.

The stable lesson is therefore smaller than “never show the reviewer the first answer.”

It is that reviewer independence is partly a property of the evidence interface.

Storage and presentation do not have to be the same system

That distinction matters for agent workflows because modern systems are getting better at preserving everything.

A work record may retain prompts, tool calls, intermediate plans, diffs, tests, model versions, review comments, and previous diagnoses. That is valuable for debugging and accountability. If a later reviewer needs to reconstruct what happened, deleting prior interpretations from the record would be perverse.

But preservation does not require simultaneous disclosure.

A review system can store the complete history while controlling the order in which it is presented. The first pass might expose repository state, test output, tool evidence, and observational trace. The reviewer records a diagnosis and confidence. Only then does the system reveal the upstream agent’s diagnosis and allow the reviewer to revise.

That design is a hypothesis, not a result of the current paper. Arslan’s experiment deletes the conclusion; it does not test a commit-then-reveal workflow on real code. The obvious next experiment is therefore about timing.

Hold the code, tests, reviewer model, and underlying trace constant. Randomize whether the reviewer sees the prior diagnosis before or after committing to a first-pass answer. Include deliberately wrong diagnoses, reliably correct diagnoses, and no diagnosis. Then measure both localization accuracy and downstream repair success.

The important outcome would not be disagreement for its own sake. A reviewer that refuses good advice merely to look independent is not better. What matters is whether the workflow can keep the benefit of correct upstream information without letting an incorrect upstream hypothesis become the boundary of the search.

Independence has more than one dimension

Software teams already know that two reviewers are not independent if they share the same failing test, the same blind spot, or the same incorrect specification.

Agent systems add another way for dependence to travel.

Two models can have different weights, different prompts, and different organizational roles while still sharing the same presented conclusion. The second reviewer may be computationally separate yet epistemically downstream.

That does not mean every review needs a blank slate. Review is usually cumulative because cumulative knowledge is useful. The engineering question is which parts of that accumulated record should be treated as evidence and which should be treated as advice.

The new paper makes that boundary visible with one short field.

Remove the prior accusation, and a reviewer that had been repeating it often starts looking elsewhere. Restore a correct accusation, and sometimes the reviewer benefits from having it.

The practical consequence is not to preserve less history. It is to stop assuming that the archive and the reviewer’s first screen should be identical. A system can keep the full record and still give the second opinion a genuine first pass at the evidence.

Sources