Skip to content

Mistakes That Never Reach the Teacher

When an imperfect checker determines which failures receive corrective attention, its false accepts are systematically underrepresented in the evidence used to improve the system unless accepted work has an independent path back into review.

One of the stranger numbers in a new verifier experiment is not 32 percent. It is 3 percent.

Dushyant Rajput set up a cost-saving language-model cascade on grade-school math. A relatively cheap Qwen2.5-7B model answered first. A second model checked its answer. Rejected answers were corrected by a teacher and then used to fine-tune the cheap model, with the hope that expensive supervision would become less necessary over time.

The loop did not improve. It degraded. On a single training seed, the independently measured error rate started around 14 percent and eventually reached 32 percent. A frozen copy of the student stayed near 13 percent.

Yet the error rate estimated through the operational checker hovered around 3 percent.

The dashboard was almost motionless while the model underneath it was getting worse.

That result comes from a single-author arXiv v1, submitted September 1 by Dushyant Rajput of AltSlate Labs LLP, and it has not yet been independently replicated. More importantly, the experiment did not demonstrate the clean self-improvement story its design was meant to test. Fine-tuning on the rejected tail destabilized the student. The paper’s long-run claim about hidden residual errors is instead supported by a theoretical model and a synthetic experiment with much simpler assumptions.

So the useful question is not whether Rajput proved that improving AI systems inevitably learn to fool their judges. He did not.

The useful question is why the 3 percent dashboard could remain so calm.

The checker had a second job

A quality checker appears to sit at the end of a workflow. Work is produced; then a test, judge, reviewer, or stronger model decides whether it can pass.

But Rajput’s loop gave the checker another role. Its verdict also determined which mistakes became future training data.

A rejected wrong answer had a rich future. It was identified as wrong, corrected by a teacher, and placed in the material used to update the student.

An accepted wrong answer had the same future as a correct one. It passed through the gate. Unless some other process inspected it, the system never learned that there had been a mistake to correct.

That changes what verification is doing. The checker is not merely measuring the worker. It is selecting the worker’s future evidence.

Rajput calls the fraction of wrong student answers that the verifier accepts the verifier’s blind spot. The name is useful because its denominator is easy to miss: it is not the fraction of all answers that pass incorrectly. It is the fraction of the answers that are actually wrong that the verifier nevertheless approves.

In a static GSM8K sweep with gpt-4o-mini held fixed as verifier, that blind-spot rate was 0.12 for a 0.5-billion-parameter Qwen2.5 student and reached 0.55 at 14 billion parameters. The three largest students, from 7B through 32B, were statistically consistent with a plateau rather than a continued rise. Because the calculation uses only wrong answers, its denominator was roughly 70 to 170 cases depending on model size, and the confidence intervals were correspondingly wide.

The verifier’s misses were therefore not a fixed property of the verifier alone. They depended on what it was judging.

On harder MATH problems, Rajput found the opposite lever. A stronger verifier could shrink the blind spot. In one Qwen2.5-7B experiment, gpt-4.1 pushed it to about 0.05. But that verifier escalated 46 percent of 200 queries while the student’s true error rate was 39 percent. In that particular setup, buying more reliable detection bought back much of the expensive review the cascade was meant to avoid.

The trade-off is not a universal price curve. It is a warning about the shape of the problem: the economic value of selective checking depends on the very errors selective checking fails to recognize.

A missing label can be created by a decision

There is an older field that makes the second job easier to see.

Researchers studying selective labels ask what happens when a decision determines which outcomes later become available for learning. In pretrial decisions, for example, the outcome under release can be observed for a person who is released, but not for someone who is detained. Maria De-Arteaga, Artur Dubrawski, and Alexandra Chouldechova described the resulting data as a form of “partial blindness”: the historical decision process helps determine which labels exist in the training record.

The analogy to software and AI review is deliberately imperfect. A patch that passes a test can still be inspected later. An accepted answer is not permanently unknowable. The missing label is usually recoverable.

But it will not recover itself.

If a coding agent turns failed tests into repair memories, the test suite decides which defects are eligible to become memories. If a support system retrains on escalated conversations, the escalation policy influences what the retraining set calls difficult. If an automated reviewer creates new guidance only after rejection, its own false accepts are absent from that guidance unless another process finds them.

The gate helps create the dataset that will later be used to improve the worker.

This is different from the simpler fact that checkers make mistakes. Benedikt Stroebl, Sayash Kapoor, and Arvind Narayanan showed that incomplete coding tests create a false-positive floor: resampling more solutions does not make an incomplete verifier complete. Eventually a wrong solution can pass.

Rajput’s loop adds a second consequence. Once rejection determines correction, a false positive can be both a missed measurement and a missing lesson.

The experiment did not prove the strongest story

That distinction also reveals the essay’s strongest rival explanation.

Rajput’s real student did not quietly improve until only verifier-invisible mistakes remained. It simply got worse. Training on the rejected tail concentrated unusually hard and style-shifted examples, and the paper reports output-format failures and collapse. Hard-tail distribution shift, label quality, catastrophic forgetting, or fine-tuning instability could explain much of that deterioration without any special conservation of blind spots.

The paper’s cleaner residual-error result comes from a synthetic 10-class model with a fixed blind region and a stationary query distribution. Under those assumptions, verifier-undetected errors can survive while detectable ones are repeatedly corrected. That is a mechanism demonstration, not evidence that production LLM systems generally converge that way.

A separate line of research also rules out a stronger claim. In the peer-reviewed weak-to-strong generalization experiments, stronger pretrained models trained on labels from weaker models consistently outperformed those weak supervisors on the studied tasks. They did not recover everything available under strong supervision, but weak labels did not impose a universal hard ceiling.

The checker can bias what the learner sees without completely determining what the learner can know.

That boundary matters. Otherwise a useful sampling problem turns into an overly grand theory of epistemic captivity.

The mistake that came back

The best independent test of the mechanism comes from an experiment designed for a different purpose.

In prover-verifier games, Jan Hendrik Kirchner, Yining Chen, and colleagues trained a strong “sneaky” prover to produce incorrect math solutions that could fool a weaker verifier. The prover found convincing mistakes. If that were the end of the loop, those false accepts would look like successes.

But the researchers had ground truth outside the verifier. They found the accepted mistakes and deliberately added them to the verifier’s next training round.

The same exploit stopped working.

The prover had to find a new one.

This is the important counterexample to any claim that verifier blind spots are inherently conserved. They are not. An accepted error can become corrective evidence if there is an independent route by which the system discovers it.

The independence is doing two different jobs at once. It tells you how wrong the operational gate is on cases the gate approved. And, if those discoveries are fed back, it changes which mistakes are available for learning.

That is why a gate and an audit are not interchangeable even when both are called “evaluation.”

A gate asks whether this case should pass now. An audit asks how often the gate is wrong, including among the cases it passed.

What to sample when everything looks fine

The practical consequence is not that every output needs frontier-model review. That would destroy much of the reason to use cascades, automated tests, or selective human oversight.

It is that a system which learns from its own review process cannot estimate the reviewer’s false accepts from the reviewer’s pass/fail judgments alone.

Some accepted work has to cross a different measurement path.

That path might be a random audit of accepted outputs, a hidden test suite that does not participate in ordinary repair, a second evaluator with genuinely different evidence, or periodic human inspection. The design can vary. The statistical job cannot: something must sample where the operational checker has already said there is nothing to see.

A falling rejection rate is otherwise ambiguous. The worker may be improving. The checker may be getting more permissive. Or the remaining errors may be shifting toward cases the checker does not recognize. A detailed database of rejected failures can still be a detailed database of only one side of the failure distribution.

Rajput could see the discrepancy cheaply because math supplied an answer key that the verifier was deliberately denied. Most consequential software and research work does not come with such a convenient oracle. The missing requirement may be the thing nobody encoded into the tests. The bad citation may be the claim nobody independently opened. The integration defect may sit beyond the boundary of the checker that approved the patch.

That makes independent auditing inconvenient precisely where it is most informative.

In the prover-verifier game, the breakthrough was not a verifier that stopped making mistakes. It was a mistake the verifier had already accepted, discovered somewhere else, and sent back.

Only then could it become a lesson.

Sources