Skip to content

When Test Writers Choose Easier Questions

For generated input-output tests, correctness is partly a property of the questions the generator chooses to face. A checker can improve its oracle score by shifting toward easier, less fault-exposing inputs, while downstream repair depends on whether the resulting tests actually constrain the right behavior.

A coding agent was solving 61.2 percent of a benchmark’s repository tasks without any generated-test feedback. Then researchers gave it tests written by another model.

Its score fell to 57.3 percent.

They changed the test-writing model and ran the same repair condition again. The score rose to 65.3 percent. With privileged oracle feedback, it reached 69.4 percent.

Those are means across three repair runs in ExecCritic, a September 2026 preprint on SWE-bench Verified. The repair policy was held fixed. What changed was the source of the feedback.

A test is supposed to make uncertainty smaller. Here, some tests made the worker worse than if it had received no generated-test feedback at all.

That result is easy to file under a familiar heading: bad tests are bad. But the more interesting question is what makes a generated test bad in the first place. A second experiment, published days earlier, exposes a decision that usually disappears behind the final green or red mark.

Before a test can judge an answer, it has to choose the question.

The hidden choice before the oracle

An ordinary input-output test contains two commitments. First, it selects an input: the situation in which the program will be examined. Then it supplies an expected output: what the program should do in that situation.

Those jobs can fail independently.

An input may be perfectly legitimate and still tell us almost nothing. If both the correct program and several plausible wrong programs produce the same result, the test is easy to pass because it does not force them apart. A more useful input reaches behavior where those implementations disagree.

Then comes the second problem. Once the test has found a revealing situation, it still needs the right expected answer. Software testing has called this the test oracle problem for decades: observing behavior is not the same as knowing whether that behavior is correct.

Human testing practice already distinguishes these dimensions. What changes when a language model generates the whole test is that one policy controls both. It chooses the input and predicts the answer.

That gives the optimizer a choice of its own. If we reward the model for producing correct expected outputs, it can get better at answering difficult questions. Or it can improve by choosing questions whose answers are easier.

Yunhao Liang and colleagues designed Correct Tests Are Not Enough to tell those possibilities apart.

They gave Qwen3.5-9B natural-language programming specifications and asked it to generate five input-output tests at a time. Audited reference programs supplied the expected outputs. A separate bank of wrong programs let the researchers ask a different question before considering the generated oracle: would this input make any of those wrong programs behave differently from the reference programs?

They called that measure potential input kill. It is deliberately bounded. Each retained task had only five to nineteen wrong programs in its bank. A high score means an input distinguishes faults represented there, not that it is a universal detector of production bugs.

The first results looked reassuring. One jointly rewarded training regime raised full-test correctness from 28.59 to 42.54 percent. Potential input kill also edged upward, from 24.06 to 25.27 percent.

Then the researchers removed the reward for fault exposure.

At the same 50-step training budget, the correctness-focused model reached 47.04 percent full correctness, higher than the jointly rewarded model’s 40.66 percent. But its potential input-kill rate fell to 21.60 percent, compared with 24.43 percent for the joint objective.

The checker had become more correct while selecting a less discriminating set of questions.

That still did not prove why. Perhaps the kill-aware model had simply become worse at computing expected outputs. To separate question selection from answer generation, the researchers performed a crossover.

Freeze the question

On 64 tasks from the training pool, each trained policy first chose its own inputs. The researchers then discarded the generated outputs. They froze the input bytes and asked the different policies to fill in the expected answers for the exact same questions.

The contrast changed the explanation.

Moving from inputs selected by the correctness-focused policy to inputs selected by a kill-aware policy increased potential fault exposure by about two percentage points. Across the answer-generating policies, those inputs imposed an average 3.81-point penalty in oracle accuracy. They were, in this measured sense, harder questions.

But when the input itself was held fixed, the kill-aware policy was not generally worse at answering it. In the main comparison, it was about 1.18 points more accurate than the correctness-focused policy on the shared inputs.

The lost oracle accuracy had followed the questions selected, not a general deterioration in answer quality.

This is the important mechanism. A correctness score for a joint test generator does not describe only how well the model answers. It is also shaped by what the model elects to be asked.

The study’s boundaries matter. The crossover used training-pool tasks, not a fresh held-out set. The main experiments used one model size and three seeds. The 142-task evaluation set had been inspected during earlier experiments, and one seed’s learning curve informed the reported checkpoint. At the best reported checkpoint, only 7.28 percent of evaluation tasks had all five generated tests fully correct.

There is another wrinkle that prevents the result from becoming a neat slogan. At the matched 50-step checkpoint, the correctness-focused model’s effective full-kill score was slightly higher than the joint model’s: 14.36 versus 13.67 percent. It chose fewer potentially discriminating inputs, but it converted more of the opportunities it did find into correct complete tests.

So “more correct tests are less useful” is not the finding. Correctness, diagnosticity, and effective detection can move separately. An objective can improve one by changing the distribution underneath another.

The tradeoff can disappear

A 2025 result makes the boundary even clearer.

Archiki Prasad and colleagues built UTGen, a system for generating unit tests to help language models debug code. They measured two properties separately: attack rate, whether an input exposed the fault, and output accuracy, whether the expected result was correct.

For Qwen2.5-7B on their harder MBPP+Fix set, random test generation achieved a 26.27 percent attack rate and 57.25 percent output accuracy. Prompting the model specifically to find failing tests pushed attack rate to 39.80 percent—but output accuracy fell to 40.78 percent.

Sharper questions were harder to answer.

Then UTGen improved both. Its trained Qwen2.5-7B reached a 41.96 percent attack rate and 58.04 percent output accuracy. A larger Qwen model moved both measures further outward.

That counterexample is as important as the new anomaly. There is no conservation law saying that a checker must choose between asking diagnostic questions and answering them correctly. Better training or more capable models can improve both.

The failure mode appears when our objective sees only part of the job. If a score rewards oracle correctness but does not care enough about which inputs are selected, the model has somewhere else to hide the cost.

In other words, the metric can improve because the model got better—or because the exam got easier.

When a measurement starts steering the worker

This distinction matters more once tests stop being passive evidence and become part of an agent loop.

In ExecCritic, a Test agent generated repository-native tests. A fail-closed harness qualified them and froze the accepted bundle. The Repair agent could change source code but not the tests. The official SWE-bench evaluator remained outside the loop as the final judge.

Under that protocol, the same base Repair condition resolved 61.2 percent of tasks with no generated-test feedback. Feedback from the base test source lowered the mean to 57.3 percent. GPT-5.6-sol-generated tests raised it to 65.3 percent. Privileged oracle fail-to-pass feedback reached 69.4 percent.

The experiment does not show that the weaker tests hurt because they chose easy inputs. Their failures could have several causes: a wrong expected result, an incomplete behavioral target, a missing edge case, or a misconception shared by the test generator and the repair process. The paper is a new preprint from a Microsoft Research-led team using a particular scaffold and model family, not a prevalence study of coding-agent failures.

What it does show is operationally simpler: feedback is not automatically helpful. A check that enters the worker’s control loop can change the sign of the intervention.

That is different from the way we often talk about tests. In a conventional description, a test evaluates work that already exists. In an agent loop, the failure becomes input to the next action. The worker sees it, forms a new hypothesis, edits a file, and runs again. The checker is now helping decide where computation goes next.

A weak check can waste effort. A wrong one can direct effort toward satisfying the wrong local target.

The distinction between measurement and control has collapsed.

A better proxy is still a proxy

It would be convenient to replace oracle correctness with mutation or fault-kill scores and declare the problem solved. Testing research gives us a reason not to.

At ICSE 2018, Mike Papadakis and colleagues studied the relationship between mutation scores and real-fault detection. Much of the raw association weakened after controlling for test-suite size: larger suites both killed more mutants and found more real faults. Yet among suites of the same size, higher mutation scores still gave significantly better guidance for real-fault detection.

Mutation-like measures can therefore be useful without being ground truth.

The same warning applies to the September experiment. Its wrong-program bank is a controlled instrument. It tells us whether a generated input distinguishes the particular faulty programs in that bank. A production failure outside the bank may remain invisible. An input may also distinguish a bank member for a reason that does not matter to a user.

That leaves a checker with several jobs that resist compression into one number. Did the test execute? Was its expected result trustworthy? Did the chosen input separate plausible wrong behavior from intended behavior? Was the checker independent enough not to share the worker’s mistake? And when the worker acted on the feedback, did the final artifact improve against evidence outside the loop?

A system can be strong on one of those questions and weak on another.

Before PASS or FAIL

A test runner ends with a reassuringly simple verdict: PASS or FAIL.

By the time that word appears, however, the checker has already made its most consequential choices. It selected a situation worth probing. It decided what should happen there. And, in an agent loop, its verdict may determine what the worker tries next.

The fixed-input crossover makes the first choice visible. The correctness-focused generator improved its answer score partly by selecting questions that were easier to answer and less likely to expose the known faulty programs. UTGen shows that this tension can be overcome rather than accepted. ExecCritic shows why the distinction matters downstream: generated tests can help a repair agent, or leave it worse than if those tests had never been supplied.

A green check tells us that the worker answered the question it was given.

It does not tell us whether the checker asked the question that mattered.

Sources