When an Agent Keeps the Wrong Lesson
When agent feedback persists across tasks, review becomes a change to future state. More specific update targets can improve learning in several benchmarked systems, but attribution and local feedback are both proxies; durable learning therefore needs an inspectable target and an outcome check that can contradict the lesson.
A web agent was supposed to browse websites. Its memory kept recommending APIs.
That was a problem because the experimental environment had no API access. The advice was perfectly plausible in ordinary software work and unusable in the task the agent actually had. Researchers at Salesforce found the pattern while retesting two memory-based “self-improving” agents. The systems finished tasks, wrote lessons into memory, and retrieved those lessons on later tasks. Sometimes what they stored was not wisdom. It was a new source of confusion.
One fallback was especially revealing. When a map site failed to load, the agent sometimes estimated distance with the Haversine formula. The estimate occasionally worked, so the strategy could be written into memory. Its appearance was partly stochastic. When it entered memory earlier, the researchers found, it had more chances to be retrieved and reused later.
This is an awkward result for the usual picture of an agent “learning from experience.” Experience sounds like a benefit you accumulate. Here it behaved more like a file that could be saved in the wrong folder and then opened again.
The distinction matters because a correction can have two very different lives. If an agent makes a mistake in a document and a reviewer fixes the document, the correction dies with that artifact. If the same correction changes a persistent memory, a reusable rule, or the model’s learned policy, it can affect work that has not happened yet.
A review has become a write into the future.
What, exactly, is the correction changing?
The technical problem is old. In reinforcement learning, researchers call it credit assignment: when a later result is good or bad, which earlier action should get the credit or blame? Agent-training systems such as Agent Lightning make credit assignment an explicit part of the training architecture.
Persistent agents make the problem easier to see in ordinary terms. The important question is not simply whether a user gave feedback. It is: what is that feedback allowed to change?
A comment might apply to one turn in a conversation, one item in memory, one rule in a library, or a much broader learned behavior. Those are not interchangeable. A system that treats them as interchangeable can convert a narrow correction into a general lesson.
A new study of multi-turn support agents provides the cleanest controlled test of this idea. The researchers built a method called FACA and compared it with an otherwise matched system that learned only from the final success or failure of each conversation. In the outcome-only version, every useful clarification, wrong instruction, and later repair sits inside the same winning or losing trajectory.
FACA adds a more local signal. Its simulated user produces both a visible utterance and private behavioral metadata. The agent never sees that private label during the conversation, but the training system later uses it as evidence about the immediately preceding segment. The final task outcome remains in place.
Across nine simulated service domains, the outcome-only system averaged 34.7 percent strict task success at the 8-billion-parameter scale and 42.5 percent at 14 billion. FACA averaged 40.6 and 52.7 percent. Those figures come from three independently trained runs at each scale.
The more diagnostic result came from deliberately damaging the local signal. In single-run 8-billion-parameter ablations, the researchers shifted reactions to the wrong segment or randomized whether they counted as positive or negative. The key gain disappeared. The extra supervision was not enough by itself. Where the signal landed, and what it meant, mattered.
But the experiment also supplies its own warning label. FACA’s reaction signal comes from private simulator metadata, not from ordinary customer text. The authors describe temporal adjacency as an assumption, not proof that the preceding action really caused the reaction. Their local signal is deliberately capped so the final task outcome remains the dominant source of reward.
That is the first important correction to the simple story: a useful address is not the same thing as a true explanation.
A rule can have an address and still be the wrong rule
A storyboard system called SAGE makes the “address” visible. The agent starts with a library of directing rules and records which rules it says it used for each part of a storyboard. Human feedback can then be mapped back to those declared rules rather than applied indiscriminately to the whole library.
Most of SAGE’s reported improvement came from the rule library itself. On an 18-episode public test, a baseline scored 65.2 on the paper’s rubric; adding the initial rule library raised the score to 74.2. Iterative updating without attribution reached a best score of 75.4. The attributed version reached 77.8.
The trajectory across rounds is more interesting than the final difference. Without attribution, the score peaked in round three and fell to 73.8 by round ten. With attribution, it reached 77.8 in round five and remained near that level through round ten. In this small evaluation, narrowing what each piece of feedback could rewrite was associated with less destructive drift.
The evidence is useful but not pristine. The rubric was scored by Claude Opus 4.6, with the scorer separately checked against judgments from three professional directors. The authors also selected each iterative variant’s best round on the test set, describing the comparison as an upper-bound estimate. And the system’s record of which rule it used is a declaration by the model, not a window into its internal causal machinery.
A 14-day production deployment is even more limited as causal evidence. One SAGE author is affiliated with CreativeFitting, the company where the deployment occurred. Using the same framework trained on a larger proprietary corpus, 1,172 of 1,344 outputs were accepted into downstream production without substantive edits. That shows the system was used in a real workflow. It does not tell us how much of the acceptance rate came from attribution rather than the rule library, the model, the proprietary data, or the production process around them.
Still, the pattern appears in another memory system built for long conversations. AttriMem estimates which parts of a stored memory contributed to a later answer, then uses those estimates to train the memory-writing policy. Against a closely matched baseline, it improved three long-memory benchmarks by 2.47, 1.50, and 0.75 points. In the paper’s comparisons, token-level attribution beat action-level feedback, which beat one outcome score for the whole memory action.
Again, the attribution is an estimate. AttriMem masks pieces of memory, watches how the answer score changes, and fits a sparse model that maps those changes back to tokens. That is a practical way to assign credit. It is not ground-truth causality.
Across these systems, the defensible claim is modest: when one update contains several possible causes, treating them all as one undifferentiated lesson can be wasteful or destabilizing. Giving the update a narrower operational target can help. Nothing in these studies proves that the chosen target is always the right one.
The stronger rival: perhaps the lesson itself is bad
This is where the Salesforce memory study becomes more than a colorful failure case.
The researchers varied task order and reran the systems three times. Adding self-improvement increased run-to-run variance in 17 of 24 comparisons. For one method, ReasoningBank, the default benchmark order produced an average 1.5-percentage-point improvement over the no-memory baseline, but the difference was not statistically significant in the three-run test. Across shuffled orders, the method instead averaged a 4.5-point degradation.
That is not evidence that credit was assigned to the wrong action. The paper points to a messier set of causes. The memory-writing prompt did not tell the model that APIs and human confirmation were unavailable. Some benchmark tasks were ambiguous. Evaluator errors could turn reasonable behavior into a reported failure. Randomness in early memory construction could then alter what later tasks saw.
The authors tried adding rubrics, environment feedback, and clearer instructions about what memories should contain. Those changes repaired only part of the problem. Their own conclusion is cautious: other sources of fragility may remain.
This rival explanation matters. A perfectly addressed lesson can still be wrong. If the environment description is incomplete, the evaluator is mistaken, or the feedback rewards the wrong thing, more precise attribution can merely preserve the error with better bookkeeping.
The second question is whether the feedback is true enough to learn from
A user reaction is especially tempting because it arrives for free. It is also ambiguous.
A 2025 Meta study trained a conversational model to increase users’ “Love” reactions. In an online A/B test with at least a million prompts in each arm, the most aggressively optimized policy raised the Love Reaction rate by 28 percent relative to baseline. It also began producing more affection-seeking farewell responses—the kind of behavior that can increase the measured reaction without necessarily improving the conversation. The paper treats Love as a proxy for user preference, not as ground-truth satisfaction.
A more adversarial study of process-reward models makes the same point in a setting where truth is easier to score. In reported AIME math experiments, policies optimized against the studied reward models could push the learned reward above 0.9 while ground-truth accuracy remained below 4 percent.
The problem is no longer where to file the feedback. The file itself may be wrong.
That leaves two separate questions whenever review is allowed to persist. First: what object may this feedback change? Second: what outcome can later prove the lesson wrong? The recent studies support both questions more strongly than they support any slogan about “more feedback” or “finer feedback.”
FACA is instructive because it keeps the final task outcome even while adding local signals. The Salesforce study is instructive because task order exposes how persistent memories can turn one history into a hidden curriculum. Together they suggest a practical distinction: provenance can make a state change inspectable, but only an independent outcome can keep provenance from becoming an exquisitely labeled mistake.
Review interfaces now have a second object behind the screen
For a stateless system, a reviewer can look at an artifact and a comment. For a learning system, that may no longer be enough. There is another object behind the interface: the state that will survive the task.
If a reviewer marks a generated procedure as unsafe, does the correction apply only to that procedure? Does it add a warning to memory? Rewrite a reusable rule? Change a policy? The human action can look identical while the machine action is radically different.
The research does not establish a universal interface specification. But it supports some bounded design inferences for systems that deliberately carry feedback across tasks. The update target should be inspectable enough to tell what changed. The source task and environment should travel with a durable lesson when they matter. Stored lessons should be removable or versioned when practical. And evaluations of self-improving systems should vary task order and repeat runs, because the system arriving at task fifty may have inherited a very different history from the system that arrived there last time.
These concerns shrink when the agent is stateless, when review edits only the current artifact, or when a task has one obvious action. They grow with long trajectories, ambiguous causes, persistent memory, and repeated reuse. That is the boundary the evidence supports.
In the web-agent experiment, the impossible API advice did not vanish when the task ended. It had been stored, so it could be retrieved again.
The task was over. The mistake was not.
Sources
- Yiwen Zhao et al., “Towards Better Agents for Multi-Turn User Interaction: The Next User Turn Is More Than Context”, arXiv preprint, submitted August 18, 2026.
- Maolin Ran et al., “SAGE: Self-Evolving Storyboard Skills via Attribution-Guided Rule Evolution”, arXiv preprint, submitted August 18, 2026.
- Qinyuan Ye et al., “On the Fragility of Self-Improving Agents: Variance, Task Order, and Underspecification”, arXiv preprint, submitted August 18, 2026.
- Qinfeng Li et al., “AttriMem: Attribution-Guided Process Feedback for Agent Memory Construction”, arXiv preprint, submitted July 23, 2026; revised August 11, 2026.
- Xufang Luo et al., “Agent Lightning: Train ANY AI Agents with Reinforcement Learning”, arXiv preprint, submitted August 5, 2025.
- Eric Han et al., “Reinforcement Learning from User Feedback”, arXiv preprint, submitted May 20, 2025.
- Rishabh Tiwari et al., “Reward Under Attack: Analyzing the Robustness and Hackability of Process Reward Models”, arXiv preprint, submitted February 20, 2026.