Skip to content

Before You Fix the Bug

Before a proactive bug-finding agent has a justified repair task, its observed discrepancy must be checked against the intended contract; a reproducible failure, plausible patch, or higher maintenance score does not independently establish that justification.

An AI agent tested a NumPy function that draws random samples from the Wald distribution. That distribution allows only positive values.

The check failed.

In August 2025, Muhammad Maaz submitted a repair for that behavior. The agent had proposed and written the failing property. Maaz had spent hours validating the result, learning the relevant code and diagnosing the numerical problem. He wrote the fix himself.

The distinction mattered in review. A maintainer confirmed the defect but needed a stronger test case to reproduce it on his own build. Later review identified an edge case in the proposed repair. The revised patch was merged in October.

The agent had found something real. It had also started work that a failing assertion could not finish: establishing a defect that the project should accept, explaining it and producing a repair that survived review.

A second report from the same research program looked similarly persuasive. This time, the software appeared to have placed Easter on the wrong day of the week.

It did not lead to a fix.

The calendar inside the complaint

Liam DeVoe’s dateutil issue began with an apparently reasonable expectation: whichever Easter-calculation method was used, the result should fall on Sunday.

For one method and the year 2015, the library returned March 30. Asking the returned date object for its weekday made the assertion fail. Another reference gave April 12.

The maintainer, Paul Ganssle, explained that two calendar representations had been combined. Method 1 returned the nominal year, month and day in the Julian calendar. But the Python date object’s weekday methods interpreted those numbers as Gregorian. In 2015, March 30 in the Julian calendar and April 12 in the Gregorian calendar represented the same Sunday. A different dateutil method performed that conversion explicitly, a distinction also present in the API documentation.

DeVoe accepted the explanation. The issue was closed as intended behavior, though Ganssle acknowledged that better support for multiple calendars would be preferable.

No one had made the failing output disappear. What changed was the obligation being assigned to it.

That is the consequential step before repair. An observed discrepancy becomes a justified bug claim only in relation to some intended behavior: a mathematical definition, an API promise, documented compatibility or a requirement from the people who use the system. Here, “returns a date” did not mean that every operation on the returned representation would respect the calendar in which its components were expressed.

The interface could still be confusing. Documentation or a different date representation might deserve work. But those would be different proposals from changing the Easter calculation to satisfy the original assertion.

Closing a computational-bug report therefore need not mean declaring a design perfect. It can mean discovering that the proposed task was the wrong response to a real source of confusion.

Before the patch has a target

Both cases came from Agentic Property-Based Testing, a project that asked language-model agents to inspect Python modules, infer properties from code and documentation, and search for counterexamples.

Property-based testing starts with a general claim about behavior rather than one handpicked input and expected answer. A test framework then tries many inputs to find a case that violates the claim. This can expose failures nobody thought to write down individually. It also leaves a crucial question open: was the proposed property actually a promise of the program?

The NumPy and dateutil records give different answers. They are selected cases from one research program, not a measurement of how often agent reports are right. Nor do they divide the work neatly into machine execution and human judgment. The agent proposed the useful NumPy property. People then validated, diagnosed, repaired and reviewed it. In dateutil, the initial human reporter also found the mistaken expectation plausible.

What mattered was the evidence that could settle the claim, not whether a person or model had written the first version.

This is the useful distinction between detecting a discrepancy and admitting an issue for action. Detection says that an observation conflicts with an expectation. Admission asks whether that expectation applies here and whether the proposed change follows from the conflict.

A stronger reproducer answers the first question more convincingly. It does not automatically answer the second. Running the dateutil assertion a million times would not decide which calendar semantics the API intended.

By the time a coding agent receives a well-supported issue, some of that work has already happened. Someone has selected a behavior, stated an expectation and supplied enough evidence to direct investigation. The ticket may still be incomplete or wrong. It rarely includes a finished diagnosis. But it gives the worker a specific claim to investigate rather than an unrestricted instruction to improve the project.

That is the starting contract of SWE-bench: a repository and an issue description, followed by an attempt to produce a resolving patch. Measuring that task is useful. The result simply cannot tell us whether the same system would have independently proposed and justified the issue it was handed.

A score for making changes

A September 2026 preprint, SWE-Prometheus, moves toward the broader assignment. Its public task asks agents to improve a repository across six engineering-readiness dimensions while preserving behavior, without supplying a concrete issue list.

Its authors also tried a revealing control: apply a fixed maintenance template without reading the repository. Across ten repositories, it earned an average 0.272 on their headroom-normalized improvement measure. That number measures rubric gains, not a percentage increase in productivity.

On one repository, Efficient-WAM, the template added a workflow and a test whose only assertion was assert True. The test passed; a separate Ruff linting probe still reported 117 findings. Those were linter findings, not 117 demonstrated functional bugs. The treatment is reported in the paper; its full patch and run evidence were excluded from the public task release.

The authors interpret the control cautiously. Some credit was available for standard maintenance artifacts without a repository-specific diagnosis. It did not establish that every credited change was useless or that the template delivered the same engineering value as targeted work.

A standard workflow or documentation file can genuinely help a project. Sometimes applying an established convention is exactly the job, and elaborate investigation would add little. The control matters because it changes what a higher score can certify. A system may have installed useful scaffolding without establishing which existing behavior violated an obligation.

That leaves two legitimate assignments which need different evidence. “Add our standard development setup” can be checked against that requested setup. “Find what is wrong with this project” requires a supported account of a problem before the resulting edit can be judged as its solution.

Removing the issue ticket does not remove human direction, either. SWE-Prometheus supplies the dimensions and the preservation requirement. It tests work within that scope, not autonomous selection of an organization’s priorities. The study does not compare the same repositories with and without a supplied issue, so it cannot measure how much effort issue formation adds.

The evidence is enough to question an inference from output to understanding. It is not enough to declare that finding problems is always harder than fixing them.

What the issue needs to carry

For proactive agents, this suggests a different object to review alongside the patch: the claim that makes the patch a response to a problem.

That claim needs an observed behavior, an expected obligation and evidence connecting the obligation to this use of the software. It should also remain possible to discover that the intended intervention is documentation, a new feature, a compatibility decision or no change at all.

This is not a demand for a large report before every edit. A broken import or an explicit request may make the obligation obvious. The point is to keep it inspectable when it is doing real work. If a later reviewer can see only a patch and a green test, the original justification may have vanished from the record.

Even a valid issue leaves priority unresolved. The two maintainer cases do not tell us which defect costs users more, which repair has the greatest opportunity cost or what should be done first. They establish something earlier: whether the alleged discrepancy supports the proposed kind of change.

The calendar report makes that boundary unusually clear. Its assertion could be reproduced, its output could be inspected, and its premise could still be revised. The successful result was a more accurate account of what the program returned, rather than a modified calculation.

A repair loop that kept changing the calculation until that assertion passed could turn a resolved misunderstanding into a new defect.

Sources