Skip to content

Was That Deleted Patch a Waste?

Related:Cards ↗

Deleted agent code is not a sufficient measure of wasted work: late requirements can reflect preventable elicitation failures or information discovered through interaction, so useful workflows need to distinguish how a constraint became knowable and preserve it once learned.

A new study of coding-agent sessions found a number that looks, at first, like a clean measure of waste.

When a user introduced a requirement after an agent had already begun editing, the next few edits removed or replaced substantially more of the agent’s earlier code. In the study’s primary matched comparison, 452 late-requirement events were followed by an average of 57.5 invalidated lines, versus 29.4 after comparable edits with no new requirement—a ratio of 1.96.

If generated code is production, the interpretation seems straightforward: requirements arrived too late, so work had to be thrown away.

Then the measurement runs into a problem. It can see when a requirement entered the conversation. It cannot see when the requirement entered the user’s mind.

That missing fact changes the meaning of the deleted code.

Two kinds of late

Requirements After the First Edit, posted as an arXiv v1 on September 2, starts from 3,553 eligible sessions in the SWE-chat corpus. The authors identify requirements stated after implementation began and, where repository history can be replayed reliably, trace later deletions or replacements of prior agent-authored lines.

The headline association survives several checks, but it belongs to a narrower sample than 3,553 suggests. Roughly 58 percent of eligible sessions passed the clean-start reconstruction screen. The paper calls that a convenience sample and does not claim it represents all eligible sessions. Its primary canonical set contains 921 late-requirement events from 402 reconstructible sessions and 74 repositories; 452 events enter the load-bearing matched comparison.

Even there, deletion is an imperfect proxy. Outcome windows overlap for 42.5 percent of the matched events. Selecting non-overlapping events preserves the direction, but the 104 events from sessions containing only one canonical late requirement produce a weaker, imprecise ratio of 1.22. In a separate semantic audit, 55 percent of sampled invalidations were judged requirement-driven, 31 percent incidental, and 14 percent unclear.

Those are not reasons to discard the result. They tell us what it is a result about: late requirement arrival is associated with extra code invalidation in this reconstructible corpus. It is not yet a census of how much agent code is wasted by bad specifications.

The deeper limitation is conceptual. The paper defines emergence as a relationship between texts. If a developer knew from the first turn that an API had to remain backward compatible but neglected to say so until turn fifteen, the constraint counts as late. If the developer did not realize backward compatibility mattered until the first implementation exposed a conflict, it also counts as late.

The repository history may look almost identical. Economically, the two cases are opposites.

In the first, the workflow failed to acquire information that already existed. In the second, making something may have been part of acquiring the information in the first place.

Deleted lines cannot tell you which one happened.

Sometimes the cheapest prototype is a question

There is strong evidence for the unromantic explanation: people often know more than the initial prompt contains, and better elicitation can pull some of it forward.

ClarifyGPT, published at FSE in 2024, treats ambiguity as something to detect before final code generation. It generates multiple candidate implementations, looks for behavioral disagreement, and asks targeted questions about the ambiguity. In a ten-participant evaluation, GPT-4 Pass@1 on MBPP-sanitized rose from 70.96 to 80.80 percent.

That is a controlled benchmark result, not a model of product development. The participants answered questionnaires containing unit-test input/output examples that helped establish the intended behavior, and the tasks were much smaller than a real repository change. But it demonstrates the rival mechanism plainly: some constraints that would otherwise arrive after code can be obtained by asking first.

This makes a simple rule such as “prototype early” hard to defend. If a question can reveal the same missing fact more cheaply, generating code first is not discovery. It is an expensive form of interviewing.

The reverse rule—“specify everything before editing”—has the same problem. It assumes that all useful information is already available in a form a person can state.

Design research gives us reason to doubt that assumption.

The artifact can change the question

In 2000, Masaki Suwa, John Gero, and Terry Purcell analyzed one practicing architect working through a design session. They tracked moments when the architect noticed an unintended visual or spatial feature in a sketch, then invented a new design issue or requirement. They also saw the reverse: a newly invented issue prompted another unexpected discovery.

The bidirectional relation covered 52 percent of that observed design process. The sample was one architect in one session, so it cannot tell us how common this mechanism is. What it shows is more basic: an external representation can participate in making a requirement articulable.

A 2001 protocol study of nine experienced industrial designers described the broader pattern as co-evolution of problem and solution spaces. Work on a solution changes the designer’s understanding of the problem; the changed problem produces a different solution attempt.

Software requirements research has found a related effect without coding agents. In a 2022 study, thirty analysts first interviewed a fictional customer and then explored comparable applications. Only 30 to 38 percent of the requirements produced after the interview could be fully traced to the customer’s initial ideas. The later app-store phase produced additional requirements, leading the authors to describe requirements as co-created rather than merely extracted.

None of these studies proves that an agent-written patch causes a developer to discover a requirement. A sketch, a competitor’s app, and executable code expose different kinds of information. But together they establish the mechanism that the coding-agent corpus cannot observe directly: sometimes interacting with a representation changes what the person knows about the problem.

This is not a new theory created by AI. What agents change is the cost and tempo at which software itself can become that representation.

A code attempt can now be cheap enough to serve two possible jobs. It can be an attempt to finish the work. It can also be a probe that makes consequences concrete enough to judge.

The same commit can fail at the first job and succeed at the second.

The useful distinction is provenance

This is why measuring churn alone loses something important.

Imagine two patches that are later deleted. In one, the agent ignored an acceptance criterion already stated in the task. In another, the user sees a working interaction and realizes that two individually sensible behaviors produce an unacceptable combination.

Both patches contribute deleted lines. Only one plausibly bought new information.

The information can also come from more than executable code. A clarifying question might expose it. A mockup might. A plan or API sketch might. The best workflow is not the one that prototypes most aggressively; it is the one that acquires the missing constraint at the lowest total cost.

That suggests a different classification for rework. Instead of asking only how much was thrown away?, ask where did the requirement come from?

Was it known from the start but omitted? Could it have been answered if the system had asked? Did the user’s preference actually change? Did an external condition change? Did the first artifact make a previously tacit conflict visible? Or was the “new requirement” really just correction of an agent mistake?

Those mechanisms imply different remedies. Omitted knowledge calls for better elicitation. Agent error calls for better execution or checking. Preference change calls for adaptability. Artifact-induced discovery calls for cheap probes—and for a way to retain what the probe taught.

Do not pay for the same discovery twice

A second new study shows why the last step matters.

Efficient Test-Time Adaptation through Human-AI Interaction, posted September 3, followed 30 people through 600 writing and data-visualization tasks. Its system collected not just messages but plan changes, direct edits, comments, and rubric changes, then used those interaction traces to alter agent context, weights, and user-specific evaluation criteria.

The adaptive agents improved solo-task success by 4.5 to 20.9 percent within 20-task sessions. The evolving rubrics also exposed failures that LLM-only rubrics missed. Against human-only rubrics, the result was more bounded: comparable failure coverage in writing and 17.9 percent more in data visualization, where the authors note that the human baseline came from general workers rather than the senior computer-science PhD annotators used for writing.

This is not a software-requirements experiment, and it does not establish that an artifact created any particular criterion. Its relevance is downstream. It demonstrates that signals expressed during work can be turned into state that influences later attempts.

That is the difference between useful iteration and recurring churn.

A requirement discovered during a failed patch has little lasting value if it remains only in the transcript that produced it. The next agent can approach the same boundary with no record of what was learned and buy the lesson again.

A living specification offers one place to put the discovery. GitHub’s current Spec Kit, for example, describes a Spec → Plan → Tasks → Implement flow and says to define what to build before building it, but it also includes clarification and refinement. A specification can be authoritative without pretending to have been complete at birth.

The durable form need not always be prose. A newly learned constraint might belong in a test, a schema, a type, a dependency relation, or another artifact that future work must encounter. The important step is converting an ephemeral correction into something that survives the session.

The September coding study cannot yet tell us how much rework belongs to each category. Its own authors point toward the missing experiment: compare late-arrival behavior with a structured elicitation checkpoint before implementation. A stronger test would also compare questions with non-code representations and executable probes, because the real issue is not whether code is magical. It is which representation makes the missing fact available most cheaply.

Until then, one conclusion is already safe.

Deleted code tells us that an implementation did not survive. It does not tell us whether the attempt was useless.

For that, we need the history of the requirement itself: what was knowable before the attempt, what became knowable because of it, and whether the system remembered the difference afterward.

Sources