Skip to content

When AI Reviewers Learn to Sound Alike

When evaluator-generated judgments become supervision for successor evaluators, their expressed judgments can narrow; whether that narrowing matters depends on whether it reduces the marginal valid coverage that additional reviewers contribute.

A peer-review system asks several people to read the same paper for a reason. The second reviewer is not supposed to be a photocopy of the first.

That sounds obvious until the reviewers themselves become training data.

In a new experiment, researchers built an AI reviewer from official ICLR reviews, used that reviewer to generate new reviews, and then trained successor reviewers on mixtures of official and generated judgments. Each paper contributed three reviews to the successor’s training set. In one condition, all three were official. In the next, one of the three was replaced by the predecessor model’s output. Then two. Then all three.

Nothing in that sequence required the successor to become harsher, kinder, or less fluent. Yet as the share of predecessor-generated reviews rose, the successors used a narrower range of scores and produced reviews that were more similar to one another in embedding space. The study, posted September 17, calls the pattern “scientific-judgment collapse.”

The label is more dramatic than the strongest result.

What the experiment establishes is a contraction in expressed distribution: the successors’ rating distributions became more concentrated, and their review text occupied less semantic space. It does not establish that the reviewers stopped catching important scientific problems. It does not show that the papers received worse decisions. It does not show a multi-generation slide toward uselessness. The paper studies one handoff, in one model family, using one predecessor reviewer.

That distinction matters because review quality has two different levels. One asks whether a reviewer is good. The other asks what another reviewer adds.

The second question is where the feedback loop becomes interesting.

One review can be excellent and still be redundant

The controlled experiment starts from Meta’s Llama 3.1 8B. The researchers first fine-tune an initial reviewer on official ICLR reviews from 2018 through 2023. They then have it generate reviews of 2024 papers and use those synthetic reviews in the four successor-training mixtures.

The successor models share the same initialization and training setup. The planned difference is how many of the three review slots per paper come from the predecessor rather than the official ICLR record. The authors use “official” deliberately: reviews released after ChatGPT became common may themselves contain unobserved AI assistance, so the zero-percent condition is not a claim of purely human authorship.

On a held-out set of 2,000 papers, the rating distribution narrowed quickly. The official-only successor had a rating standard deviation of 1.63 and entropy of 2.31. Replacing one of the three reviews per training paper with a predecessor-generated review reduced those measures to 1.44 and 2.14. At higher synthetic shares, both remained below the official-only baseline. The mean score, meanwhile, moved up and then back down rather than drifting steadily toward leniency or severity.

The text moved too. The average semantic distance among independently generated reviews of the same paper fell from about 0.159 in the official-only successor to 0.142 in the all-synthetic successor, roughly an 11 percent reduction. A corpus-level spread measure fell by about 5 percent over the same comparison. Both decreased monotonically across the four synthetic-exposure conditions.

Those are clean differences in the model’s outputs. They are not yet a clean measure of what a scientific review is for.

Embedding distance can capture whether two pieces of prose occupy similar semantic territory. Rating entropy can capture whether a model uses a broad or narrow portion of a score scale. Neither tells us whether the second review found a valid flaw the first one missed.

A panel of reviewers does not need maximum disagreement. Three people inventing three different mistakes would be diverse in the least useful possible way. What the panel needs is non-redundant valid coverage: important concerns that are correct, well supported, and not already supplied by everyone else.

The fresh experiment makes that quantity newly urgent without measuring it directly.

The useful unit is the criticism, not the reviewer

A separate 2026 study gets closer to the missing measurement.

Researchers asked 45 domain scientists to evaluate 2,960 individual criticisms drawn from human and AI reviews of 82 Nature-family papers. Instead of treating an entire review as good or bad, the experts judged atomic criticisms for correctness, significance, and evidentiary support.

The result makes a simple “AI reviewers are worse” story difficult to sustain. On the study’s composite measure, a GPT-5.2 reviewing agent averaged 60.0 percent, compared with 48.2 percent for each paper’s top-rated human reviewer. Across the AI systems tested, accurate AI criticisms were often significant and well evidenced. About 26 percent of AI-raised criticisms had no similar counterpart in the human reviews.

Then came the part that changes how those quality scores should be read. Criticisms from different AI reviewers were much more likely to overlap than criticisms from different human reviewers: about 21 percent versus 3 percent of cross-reviewer criticism pairs were judged similar.

Those findings can coexist because individual quality and panel value are different measurements.

Suppose the first AI reviewer finds ten excellent problems. A second reviewer that finds the same ten may also be excellent. But the second reviewer has contributed little additional coverage. A reviewer that finds eight of those ten and three different valid problems might score similarly as an individual while being much more valuable to the panel.

This is why homogeneity can matter before it becomes incompetence. The risk is not that similar reviewers must be wrong. It is that the marginal value of another review can shrink even while every reviewer still looks strong in isolation.

The Nature-family study does not establish that recursive training caused its AI-AI overlap. Those models were not produced by the feedback loop in the ICLR experiment. But it supplies an independent demonstration of the practical quantity the recursive study leaves open: how much valid information another reviewer actually adds.

Not every kind of disagreement buys coverage

Once review diversity is framed as a resource, it is tempting to optimize for difference itself. Peer-review research gives a reason not to.

A 2024 study of roughly 5,000 conference submissions examined several kinds of diversity among reviewer groups and their relationship to two outcomes: how much of a paper or its review criteria the group covered, and how redundant the reviews were.

The effects depended on what “diverse” meant. Reviewer groups that were topically diverse, mixed in seniority, or came from distinct publication networks showed broader coverage on the study’s measures. Organizational and geographic diversity did not show the same coverage increase. Several diversity dimensions were associated with lower redundancy, but again the pattern was not uniform.

That is a useful warning for AI systems. More varied wording is not automatically better checking. More varied scores are not automatically better checking. Even demographic or institutional variety is not interchangeable with the particular differences in expertise, priors, methods, and failure sensitivity that cause reviewers to notice different things.

The target is not diversity in the abstract. It is the set of valid concerns the group can recover.

This also changes the meaning of the new recursive-training result. A decline in semantic spread is a signal worth investigating because it could indicate that the reviewer population is becoming more redundant. But the consequential test is not whether embeddings move closer together. It is whether adjudicated issue coverage moves with them.

If review text becomes stylistically more alike while the set of correct, important criticisms stays constant, the apparent collapse is largely cosmetic. If rare but valid criticisms disappear as predecessor-generated supervision increases, then the narrowing has become an operational failure.

The experiment has not yet separated those possibilities.

The feedback loop is real, but its outcome is conditional

There is a reason the authors reach for the language of collapse. Machine learning already has a well-developed cautionary story about systems trained on their own descendants.

In 2024, a Nature paper on recursive synthetic-data training showed that successive generative models can lose information about the tails of the original data distribution when generated data replace the data that preceded them. Rare regions are easy to undersample. Once they vanish from the generated set, the next model has no example from which to recover them.

That mechanism makes intuitive sense for evaluators too. If a predecessor reviewer rarely produces a certain kind of objection, a successor trained heavily on that predecessor receives fewer examples of that objection. What the predecessor leaves out can become harder for the successor to express.

But the broader model-collapse literature also supplies the boundary. Collapse is not a law that fires whenever a synthetic example appears. A separate study found that, in the regimes it tested, accumulating original real data alongside successive synthetic generations avoided the collapse seen when each generation replaced the old data. Other work has likewise found that the size and direction of recursive distribution shifts depend on properties of the starting data.

That matters here because the ICLR reviewer experiment is a replacement-style intervention inside a fixed three-review slot. As synthetic exposure increases, explicit predecessor outputs take the places of official reviews. It does not test every plausible data regime for future reviewer training.

It also uses one synthetic teacher. That leaves open a different explanation for the narrowing: perhaps the crucial variable is not “synthetic” supervision in general but learning repeatedly from one teacher’s characteristic score calibration, prose style, and set of preferred concerns.

A multi-teacher condition would distinguish those stories. So would multiple independent training seeds. So would an issue-level evaluation that asks experts to mark which valid criticisms survive each handoff.

Until then, “recursive sample compression” and “single-teacher monoculture” remain partly entangled.

AI participation is not the treatment

There is another easy story the evidence rules out: that putting AI anywhere in peer review makes the process narrower or worse.

At ICLR 2025, researchers ran a randomized experiment involving more than 20,000 reviews. Some reviewers received optional feedback generated by a guarded LLM system that flagged vague comments, possible misunderstandings, and unprofessional language. Twenty-seven percent of reviewers who received the feedback changed their reviews, incorporating more than 12,000 suggestions. Among the updated reviews, the resulting text was longer and was judged more informative by blinded researchers; selected reviewers also participated more in rebuttal discussions.

That intervention is almost the opposite of the recursive-training experiment. The model did not replace a reviewer and its output was not being used to train a successor reviewer. It acted as feedback to a person who retained authorship and could accept or reject the suggestion.

The lesson is not that assistance is always safe. It is that “AI in review” is too broad a category to explain anything useful.

Assistance, autonomous review generation, panel composition, and later training on review text are different causal interventions. A system can benefit from one and still be vulnerable to another.

The same distinction applies outside science. A coding agent that proposes a review comment is not the same system as a reviewer model fine-tuned on thousands of its own earlier comments. A benchmark judge used for one evaluation is not the same thing as a judge whose verdicts become the labels for its successor. Once evaluation artifacts cross that boundary, provenance stops being clerical metadata and becomes part of the learning process.

What to measure before the loop closes

If organizations begin reusing machine-generated checks as training data, the obvious dashboard is likely to show single-review quality: agreement with a reference score, defect precision, helpfulness ratings, perhaps average acceptance accuracy.

Those are necessary measurements. They are not sufficient.

A second family of measurements asks what the reviewer population covers together. How many distinct valid issues does the second reviewer add? Which severe or rare defect classes disappear first? How much overlap comes from genuine consensus on important flaws, and how much is duplicated attention to easy ones? Does a panel built from several model families recover more issues than several copies of the same reviewer? Does retaining independent human or reference supervision preserve marginal coverage across retraining cycles?

The new ICLR experiment suggests that these questions should be asked before a reviewer becomes visibly bad. Its successors still produce reviews. Their average score does not march toward some absurd extreme. The first warning is statistical concentration.

That makes lineage important. If two reviewers were trained on substantially the same generated judgments, calling them “independent” because they run in separate processes can be misleading. Their errors may share ancestry.

But provenance alone is not a cure. A label saying “synthetic” does not tell us whether a review is redundant, wrong, or valuable. The point of lineage is to make the relevant comparisons possible: single-teacher versus multi-teacher supervision, replacement versus accumulation, shared ancestry versus genuinely independent training data.

The hardest experiment is also the one that would make the story consequential. Train the successor reviewers under those different regimes, then have experts adjudicate the atomic criticisms they produce. Count not merely how different the prose looks, but how many distinct correct and important problems each additional reviewer contributes.

The current evidence leaves open a reassuring outcome. The successors might sound more alike and use a tighter score range while preserving the full set of useful criticisms. If so, the dramatic phrase “judgment collapse” would describe a distributional change more than an epistemic one.

It also leaves open the opposite. Rare objections may be exactly the kind of tail behavior recursive supervision drops first.

A review system should be able to tell the difference.

The reason to ask several reviewers was never to maximize the number of reviews. It was to increase the chance that something one reviewer missed would be caught by another. If machine-generated judgments become the material from which future judges learn, that marginal contribution is the quantity to protect.

The second reviewer is supposed to add something.

Sources