When Every Failure Gets a Fix
A failed agent run underdetermines the repair it justifies: root cause, repair surface, and repair scope can differ, so shared harness changes should be treated as scoped hypotheses and tested across the contexts they claim to govern.
In the public code for a new agent experiment, one improvement strategy follows an almost comforting rule: collect the failed runs, then call the harness updater once for each failure. Every red mark earns its own chance to make the system better.
On the benchmark table, that responsive system lost to the harness that never changed.
Ecdysis, released September 10, compares several ways of improving the runtime machinery around language-model agents. Across ten model-domain cells in the current public results, a fixed human-configured harness averaged 51.67 percent accuracy. The serial system that evolved from individual failures averaged 46.67 percent. A different procedure that compared failures across tasks before proposing changes reached 54.67 percent, and a more elaborate version with multi-role diagnosis reached 59.33 percent.
The easy story is that one system learned from its mistakes while the other learned the wrong lessons. But a failed run contains less instruction than that story assumes. It says that the whole arrangement did not produce the required result. In an agent, that arrangement may include a model, system prompt, tool schema, memory layer, repository, external service, and grader. The red mark does not arrive with an address telling the updater which component should change.
Before a system can learn from a failure, it has to decide what kind of lesson the failure is evidence for.
The red mark has no address
Engineers already know this in ordinary debugging. An HTTP 500 means a request failed. It is not an instruction to rewrite the web server. The same symptom can come from a bad query, an expired credential, a malformed request, a downstream outage, or a bug in the server itself.
Agent systems make this ambiguity harder to ignore because more of the surrounding machinery is becoming editable. A self-improving system can change prompts, tool descriptions, orchestration code, retries, memory rules, graders, and sometimes the model or model selection policy. The existence of many available repair surfaces makes a failed run less, not more, specific about what should happen next.
This is already an explicit research problem. A July paper called Model or Harness? calls it repair assignment. The authors organize 41 agent failure modes around interactions among models, harness components, users, tools, environments, and graders. Their point is that the same system-level outcome can imply very different interventions depending on where the failure originated.
Their strongest automated judge reached a Cohen’s kappa of 0.76 against human category labels. That is evidence that the taxonomy can be applied with substantial consistency. It is not proof that a label identifies the uniquely best intervention. A diagnosis can be reproducible and still leave engineering choices open.
That distinction matters because three questions that are easy to collapse are actually different.
Root-cause location asks where the failure originated. Repair surface asks what component it is practical to change. Repair scope asks which tasks, repositories, models, tenants, or environments should inherit that change.
Suppose one model family repeatedly emits an invalid argument for an otherwise sound tool API. The model may be the root cause. The cheapest repair surface may still be the harness: add a constrained schema, a converter, or a recovery step. But the right scope may be only that model family. A model-originated weakness can rationally produce a harness-side fix without justifying a global harness rule.
Once those three addresses are separated, “learn from failure” stops being one operation.
The new experiment changes several things at once
Ecdysis is interesting because its public implementation exposes exactly how a seemingly reasonable learning loop can change shape.
The serial baseline, labeled E3 in the repository, evaluates a training set, collects failed trajectories, and then makes one harness-updater call per failure. The resulting edits are composed into the harness used for the next round.
The cross-instance condition, E4, does something different. It clusters failures by termination reason and failed reward basis, then makes one batched updater call over the grouped evidence. Failure groups that span more distinct task IDs are treated as more important evidence. The full E5 method adds a diagnosis process with analyst, critic, engineer, and moderator roles before the harness is edited.
The current public result table is striking. But it does not isolate one causal ingredient. Moving from E3 to E4 changes the evidence presented to the updater, the number and order of updater calls, the amount of context available at once, and the opportunity for sequential patches to interfere with one another. E5 changes the inference process again. The repository is also explicit that this is an ongoing research release and does not include the raw benchmark data, generated traces, local experiment configurations, or private run artifacts needed to reconstruct the published table independently.
So the result does not establish that recurrence across tasks correctly identifies which layer “owns” a failure. It establishes a narrower fact about this experiment: serial per-failure harness evolution underperformed the fixed harness, while methods that aggregated failures and used more structured diagnosis outperformed it.
That is enough to change the engineering question.
If an individual failure is weak evidence for a global repair, what makes a failure pattern strong enough to justify one?
Recurrence is evidence about scope
Ecdysis uses recurrence across task instances as one answer. A failure pattern seen on several tasks is treated as more likely to reflect something systematic in the shared harness than a pattern seen once.
That is a useful inductive bias. It is not a law.
Two task IDs may not be independent evidence. They can share the same template, reward function, API, repository structure, or benchmark convention. A grader bug can recur across many tasks because the grader is shared. A model-specific quirk can recur across many tasks because the same model is shared. A genuinely global harness defect can appear only once if the triggering condition is rare.
What recurrence can do is change the prior. If the same failure survives changes in repository, task family, and model family, a shared interface or harness defect becomes more plausible. If the pattern follows one model across otherwise different environments, a model-relative repair becomes more plausible. If it follows one repository regardless of model, a repository-local explanation deserves more weight.
The important word is follows. The pattern becomes informative when we vary the surrounding context and see what it travels with.
This is why the strongest practical interpretation of Ecdysis is not “patch only recurring failures.” It is “treat patch scope as something to infer from comparative evidence.” Recurrence is one signal among several: independence of the contexts, similarity of the failure mechanism, cross-model transfer, regression behavior, and held-out performance all matter.
Another recent system, HarnessCompass, reaches the same general problem from a different direction. Instead of relying on recurrence as the main guardrail, it constrains candidate modifications to be task-agnostic, adds direct feedback about harness usage, and optimizes components separately to reduce interference. The method reports stronger held-out and cross-model transfer than prior harness-evolution approaches. Whatever one thinks of the specific numbers, it is a useful counterexample to the idea that recurrence is the only route to better scope control.
Sometimes the worker’s limp belongs in the harness
The original Ecdysis framing distinguishes “model-specific accommodation” from “harness-level repair.” That language makes accommodation sound like contamination: the harness bends around one worker’s weakness and becomes less general.
Sometimes that is exactly what should happen.
Self-Harness, released in June, starts from the opposite premise. Different language models behave differently, so the best harness may legitimately be model-specific. The system mines recurring weakness patterns from one model’s traces, proposes small harness changes, and admits them only after regression testing.
On held-out Terminal-Bench tasks, the reported pass rates for its three tested models rose from 40.5 to 61.9 percent, 23.8 to 38.1 percent, and 42.9 to 57.1 percent. In that deployment model, compensating for a model weakness in the harness is not a failure of abstraction. It is the product being optimized.
That counterexample removes the moral coloring from “accommodation.” A patch is not wrong because it is specific. It is wrong when its declared or actual scope exceeds the evidence for which it has been validated.
A recovery instruction that helps one model and passes regression tests for that model can be a good model-local patch. The same instruction silently inserted into a shared multi-model harness may be an overgeneralization. The code could be identical. The mistake is in where it is allowed to govern.
This is a familiar idea in configuration management: a rule can be correct in one environment and harmful when promoted globally. What changes in agent systems is that the promotion decision may itself be automated.
A patch changes the evidence that comes next
There is another reason scope matters early. Harness evolution is path-dependent.
Imagine a system observes a failure, adds a global workaround, and runs the next task. That next trajectory is now evidence about a different system. If the workaround suppresses one symptom while introducing another, later failures are shaped by the earlier diagnosis. A sequence of locally plausible edits can therefore turn the harness into a history of prior assumptions.
The Ecdysis serial baseline makes this possibility visible because its updater acts once per failure and composes those changes. The public experiment does not report a causal decomposition of patch interference, so it would be an overreach to say interference explains the performance gap. But the architecture creates the opportunity for order effects that the batched condition reduces by construction.
This is where the difference between detecting failure and managing change becomes important. A self-improving runtime needs more than a queue of red traces and an editor. It needs enough provenance to know which model, harness version, repository state, task family, and evaluator produced each failure; enough comparative evidence to estimate whether a proposed change is local or shared; and enough held-out testing to decide whether the claimed scope survived contact with new cases.
In other words, the object being reviewed is not only the patch. It is the patch plus its validity domain.
A useful change record might say that a modification is intended for one task, one repository, one model family, or the global harness. That declaration is itself testable. A model-local patch can be run against other models to see whether it transfers. A supposedly global tool fix can be exercised across repositories. A rare singleton can be investigated without automatically becoming infrastructure.
The goal is not to make every failure wait for a committee of identical failures. Some one-off incidents reveal genuine global defects. The goal is to stop treating the visibility of a failure as evidence that the broadest editable layer should learn from it.
What the failure actually tells you
The interesting result in Ecdysis is not that a newer harness optimizer scored higher than an older one. It is that the apparently more responsive system—the one that changed after each failure—did worse than the fixed human harness it was meant to improve.
The paper’s method suggests one explanation: compare failures before deciding which ones deserve shared repairs. Its own public implementation also leaves strong alternatives on the table: batching may provide better context, fewer updater calls may reduce interference, and structured diagnosis may simply improve proposal quality. A clean experiment would hold evidence, compute, and edit opportunities constant while varying only the recurrence information. That experiment has not yet been reported.
But the uncertainty does not erase the design lesson. It sharpens it.
A failed agent run is evidence about an outcome. Turning that outcome into a code change requires at least two additional judgments: where to intervene and how broadly the intervention should apply. Recent failure-localization work makes the first judgment explicit. Ecdysis makes the second hard to ignore. Self-Harness shows why the two cannot be collapsed: the root cause can be model-specific while the practical repair sits in the harness, and that can be entirely correct when the harness is intentionally scoped to that model.
The next generation of self-improving agents will not be dependable merely because they notice more failures or repair them faster. They will need to preserve what each repair was supposed to be true for. When a fix is promoted from one model or repository into a shared harness, that promotion should be treated as a new claim requiring evidence at the broader scope.
The failed run supplied the first observation. It did not supply the scope.
Sources
- Ruiqing Yue et al., “Ecdysis: Efficient and Effective Training of Runtime Harnesses for LLM Agents”, arXiv v1, September 10, 2026.
- Ecdysis public implementation and current results, accessed September 11, 2026.
- Harsh Raj et al., “Model or Harness? An Interaction-Centric Taxonomy for Localizing Agent Failures”, arXiv, July 30, 2026.
- Hangfan Zhang et al., “Self-Harness: Harnesses That Improve Themselves”, arXiv, June 8, 2026.
- Luan Zhang et al., “HarnessCompass: Guiding Automatic Harness Evolution toward Generalizable and Effective Agent Harnesses”, arXiv, August 3, 2026.