Skip to content

Two Actions Left, Two Very Different Agents

In closed-loop agent work, endpoint success mixes which states a policy reaches with how well it acts from a held-fixed state; checkpoint handoffs can separate those contributions without explaining the mechanism that produced them.

The hard part of one household task was already over.

An agent in ALFWorld, a text-based benchmark for embodied tasks, had picked up a pencil from a desk. Six shelves were available. From that exact state, the environment could verify that success was two valid actions away: walk to a shelf, then put the pencil there.

Researchers made two exact copies of the state. Both continuation models received the same observation, visible history, admissible actions, remaining two-action budget, and paired randomness. One model described the right plan, then spent its first action looking around. Its second action reached a shelf. The budget ended before it could put the pencil away. The other model walked to a shelf immediately and placed the pencil on its second action.

That sequence appears in Xuan Liu and Jingbin Qian’s September paper, Reach or Solve? Attributing Agentic RL Gains with Checkpoint Handoffs. It is a small experiment with an unusual virtue: the inherited problem has been held still. Only the solver changes.

Normal agent evaluation rarely gives us that comparison.

A long-running agent changes the observations, partial work, mistakes, and information that its later actions inherit. Two policies can start with the same task and, halfway through, be working from very different states because of what they did earlier.

The first half partly writes the test the second half must take.

One score, two different jobs

The underlying idea is old. Sequential decision-making has long distinguished the states a policy visits from the value of acting well once a state is fixed. The useful contribution here is experimental: the authors make those two quantities separately observable in live language-model agents without retraining them.

They assign two roles. A reacher acts first and produces an intermediate state and history. A solver inherits that state and tries to finish. Crossing supervised-fine-tuned and reinforcement-learned checkpoints in both roles creates comparisons that an ordinary endpoint score cannot.

ALFWorld makes the separation especially concrete. The researchers defined a frontier of states that the environment could certify were exactly two valid actions from success and still had enough budget to take those actions. Across 30 unseen tasks with four paired seeds each, the supervised checkpoint reached that frontier in 13.3 percent of trajectories. The reinforcement-learned checkpoint reached it in 82.5 percent.

Then they replayed every reached state into identical clones and gave both solvers the same two-action continuation. Pooling the 115 unseen-split frontier states produced by either reacher, the supervised solver completed 56.5 percent. The RL solver completed 95.7 percent.

Those are two different advantages. The RL checkpoint gets to the verified frontier much more often, and it also finishes more often after the inherited state is fixed.

A second benchmark, TravelPlanner, used a different and later handoff. Across 180 tasks at each of three model scales, the reacher-by-solver interaction was positive at all three scales, from +10.6 to +17.8 percentage points. But that result has a narrower interpretation: the handoff occurs immediately before the reacher’s final model call, so the inherited object is mostly accumulated message and tool history rather than a compact mid-run physical state.

The engineering distinction is still useful. If a system rarely reaches a workable intermediate state, a stronger finisher may not fix the bottleneck. If two systems reach comparable states but one cannot finish from them, more exploration attacks the wrong problem.

The paper calls the two quantities Reach and Solve. In ordinary language: what did the first part leave behind, and how well could the second part use it?

The RL history was not simply an easier world

A tempting interpretation is that reinforcement learning taught the agent to prepare an objectively easier future for itself. The experiment does not show that.

On the unseen ALFWorld states produced by the supervised reacher, changing only the solver from supervised to RL raised conditional completion by 50 percentage points. On states produced by the RL reacher, the same same-state switch raised completion by 37.4 points. Once the frontier had already been reached, the RL solver’s advantage was actually larger on the supervised reacher’s states.

The positive endpoint interaction comes from a different quantity: arrival-weighted completion over the full task population. Non-arrivals count as failures. Under the supervised reacher, switching solvers raised endpoint success from 5.8 to 12.5 percent, a 6.7-point gain. Under the RL reacher, it raised success from 48.3 to 79.2 percent, a 30.8-point gain. The RL reacher supplied many more viable handoffs, so the solver upgrade affected many more trajectories.

That is a behavioral interaction, not a mechanism.

The study cannot tell us whether the RL checkpoint reaches useful states because it explores better, avoids irreversible mistakes, loops less, extracts information differently, or benefits from some other property of the training pipeline. It also cannot tell us whether an inherited history is broadly useful or especially compatible with the policy family that produced it. The two training pipelines are independently released, but every evaluated checkpoint belongs to the broad Qwen family.

Even state is broader than the simulated room. The paper’s evaluation state includes the task, environment configuration, visible interaction history, and remaining budget. Useful inheritance may live in the external world, the transcript, prior tool results, or their combination.

So the result is not evidence that RL agents have learned to “scaffold themselves.” It is evidence that an endpoint score can hide two contributions that need different interventions to separate.

The missing trajectories are part of the result

There is an obvious shortcut: if the two reachers produce different states, compare the solvers only after a reacher has arrived somewhere useful.

That changes the population being measured.

On the unseen ALFWorld sample, the supervised reacher arrived at the two-action frontier in only 16 task-and-seed trajectories drawn from six tasks. The RL reacher arrived in 99 trajectories drawn from 28 tasks. Keeping only arrivals therefore compares solver performance over two different survivor populations.

The paper includes a post-hoc diagnostic that shows how much this can matter. Keep non-arrivals as failures, and the reacher-by-solver interaction is +24.2 percentage points, with a 95 percent interval from 13.3 to 35.0. Delete non-arrivals and compute a survivor-only version, and the estimate becomes −12.6 points, with a wide interval from −45.8 to 22.7.

The negative point estimate is not evidence of a negative conditional interaction. The authors say so explicitly: the interval is wide, and they specified this diagnostic after seeing the point estimates. What it demonstrates is that deleting non-arrivals changes the estimand.

That is a familiar selection problem. Arrival depends on both the policy and task difficulty, so conditioning on arrival can select away part of the effect you are trying to explain. For population-level endpoint performance, the failures must remain failures. For the solver effect, identical cloned states supply the cleaner comparison.

The handoff creates the counterfactual an ordinary log does not contain: what would another policy have done from exactly here?

Coding agents make the same problem visible

The software version is easy to recognize. A coding model that edits the wrong file at step three does not merely make one bad move. It changes which tests fail, what diffs exist, what the model reads next, and which repairs remain cheap. A stronger model introduced at step ten inherits that history; it is not solving the state it would have created for itself.

An independent August study, The Replay Gap, tested this directly on SWE-bench trajectories. Instead of evaluating a model switch by stitching together pre-recorded logs, Ashritha Gonuguntla forked live trajectories, reconstructed the environment and message prefix, and let the substituted model continue in its own branch.

Across roughly 900 rollouts, model swaps rewrote 61 to 94 percent of post-fork actions. In early swaps, 74 to 77 percent diverged on the first post-fork action. All five observed outcome flips occurred in swap branches; none occurred in 359 same-model control forks.

That study does not perform the Reach/Solve decomposition. It establishes a complementary fact: changing the worker can rapidly change the trajectory that later work inherits, so static replay can grade states the switched model would never have produced.

This also limits the novelty story. Live branching of coding-agent trajectories already existed. Earlier work such as AgentBoard and Process Evaluation for Agentic Systems had already argued that final success can hide important intermediate behavior. The new paper’s narrower contribution is the crossed reacher-and-solver experiment: hold the inherited input fixed, swap the continuation policy, and keep non-arrivals in the population-level accounting.

A coding checkpoint is more than a Git commit

For software agents, that design suggests a useful experimental primitive. A Git commit can preserve the repository portion of an intermediate state. A trajectory record can preserve the visible plan, tool observations, model context, and remaining budget. Together they can make a handoff reproducible enough to ask questions that an endpoint benchmark cannot.

But a commit is not the whole state. Services, caches, credentials, uncommitted files, external tool effects, and other runtime details may matter. A coding handoff has to state exactly what it preserved before claiming that two solvers received the same problem.

That caveat is not bookkeeping. It defines the experiment.

With enough state captured, a team can ask: Would another model finish from this checkpoint? Did a new harness help mainly by getting to better intermediate work, or by making better decisions from the same work? Does a checkpoint transfer cleanly across solvers, or is its value tied to the policy that produced it?

Those questions can point to different engineering changes even when the top-line success rate is identical.

The pencil example makes the distinction memorable because the researchers froze the world at exactly the point that mattered. Pencil in hand. Two actions left. Same room, same history, same budget. In that state, the RL solver was plainly better.

But most supervised trajectories never reached a comparable state at all.

One score can tell you which full system won. It cannot tell you whether the advantage came from reaching the right problem, solving the same problem better, or the interaction between the two. Once an agent is allowed to change its own future inputs, answering that question requires an intervention that holds the inherited input fixed. A checkpoint handoff is one way to do it.

Sources