Right Values, Wrong Places
In controlled structured-generation tasks, a model can preserve the right components while binding them to the wrong roles; separating those failures helps identify when deterministic structure can remove mechanical work and when added structure would instead make a system more brittle.
A language model was given a synthetic billing record to complete. It did not forget the zip code. It did not invent a different city. The values it needed survived.
The record was still wrong.
The model had put correct values under the wrong JSON paths. For a reader skimming the output, that can look like an almost-correct answer. For software consuming the record, the distinction is absolute. A valid zip code in the city field does not become partly useful because the model remembered it accurately.
The example appears in an August 26 preprint called Where vs What. Its useful move is not a new theory of intelligence. It is a new way to ask what failed.
Did the model lose the value? Or did it lose the value’s role?
Two ways to be wrong
Yiwei Zhang and his colleagues constructed JSON and table tasks with deterministic answers. Because the correct values and their intended positions were known in advance, the researchers could score two things separately.
Value presence asks whether a required value appears anywhere in the output. Value placement accuracy asks whether it appears where it belongs.
That distinction produces an unusual diagnostic picture. At the hardest tested JSON level, DeepSeek-V4-Flash retained 95.6 percent of the planted values. Yet 35.4 percent of the values it recalled were misplaced. Qwen2.5-7B showed the same kind of failure more severely: its displacement rate reached 73.8 percent.
The models were not simply getting worse at everything at once. In these tasks, information could remain available after the relationship between the information and its structural address had begun to fail.
The researchers then changed the task rather than the model. Deeper recursive JSON trees increased displacement. Merely adding more planted values at a fixed complexity did not produce the same steady trend. Replacing meaningful field names with opaque labels also made placement worse. So did reusing the same field names in different branches.
Those tests make one explanation plausible: models may use semantic labels and positional cues as shortcuts for locating where a value belongs. As the number of similar-looking destinations grows, the cues become less reliable.
But the paper does not observe an internal “topology module” breaking. Its authors call the shortcut account a plausible interpretation, and their own later experiment supplies a reason not to turn the measurement into a grand cognitive claim.
They trained a 7-billion-parameter model with rewards aimed directly at structural placement. JSON placement accuracy rose from 0.264 to 0.629 and generalized well to unfamiliar JSON schemas. Yet on an out-of-distribution table task, coordinate placement stayed essentially at baseline even while table formatting improved dramatically. A skill that transferred cleanly across every kind of structure would have produced a different result.
So binding correctness is most useful first as a diagnostic category. A system can choose the right component and still attach it to the wrong role. That does not require us to believe there is one universal mental faculty responsible for every such error.
What happens when the relationships multiply
Five days after Where vs What appeared, four Microsoft researchers posted a paper about a problem that looks less like a laboratory exercise and more like enterprise software: turning natural-language contact-center rules into executable workflow graphs.
A routing rule is not just a bag of actions. “Raise priority,” “wait ten seconds,” and “end the conversation” can all be correct actions and still form the wrong workflow. Each action has to be connected to the right condition, branch, fallback, and predecessor.
The researchers built a benchmark of 635 manufactured rules in the dialect of a commercial Microsoft routing platform. The data are synthetic rather than observed production failures, and the underlying platform vocabulary and evaluation set are proprietary. The benefit of that artificiality is exact ground truth: for every rule, the correct graph is known.
On the stronger reasoning models, choosing the action types was almost solved. GPT-5.2 and Opus 4.6 identified the correct action-node set roughly 98 to 100 percent of the time across configurations. GPT-5.3-chat was weaker, ranging from 82.4 to 92.6 percent, but even there node selection was consistently easier than getting all of the node parameters exactly right.
The failure became more visible as the graph grew dense. Under the monolithic GPT-5.3-chat prompt, exact condition accuracy was roughly 93 percent for rules with one or two condition-action pairs and fell toward zero once a generated rule contained more than fifteen pairs.
Again, “the model misunderstood the request” is too coarse a diagnosis. It could identify much of the correct action vocabulary while losing the assignments among actions, predicates, and branches.
The Microsoft team changed the architecture around that observation. Instead of asking the model to write the complete target graph, they asked it for a smaller intermediate description. Deterministic code expanded the repetitive combinatorics into final JSON. Another stage selected only the relevant action vocabulary before generation.
For GPT-5.3-chat, that full pipeline raised exact condition accuracy from 56.5 to 82.2 percent and the paper’s judge-validity measure from 56.4 to 80.6 percent. It also cut total per-rule token use from 21,264 to 11,980.
That result is easy to turn into an overly convenient slogan: take structure away from the model and give it to a compiler. The paper itself is more complicated.
The intervention changed several things at once: the representation, the available vocabulary, the amount the model had to emit, and the deterministic work done afterward. Some error classes nearly disappeared, but others did not. In the authors’ automated analysis of 2,348 failing cells, wrong edge ordering was the largest category, accounting for 49.5 percent of failures. It barely improved under some partial decompositions and fell meaningfully only under the full system. Cross-variable AND/OR errors stayed roughly flat.
The point estimates have another boundary. Each evaluation cell was one temperature-zero generation, so the authors’ intervals capture variation across rules and scorers, not sensitivity to paraphrased prompts or repeated stochastic decoding.
The experiment therefore shows something narrower and more useful than “compilers fix model reasoning.” When relationships are known well enough to formalize, moving some mechanical assembly into deterministic machinery can remove work that the model was performing unreliably. It does not remove every relational error.
A valid structure can still be wrong
This is why syntactically valid JSON is such weak reassurance.
Constrained generation can guarantee that braces close and a schema is obeyed. Where vs What explicitly separates that outer validity from correct placement. The Microsoft pipeline reached 99 to 100 percent valid JSON while still making mistakes in conditions, Boolean logic, and edge ordering.
The distinction is familiar elsewhere in software. A program can compile and still do the wrong thing. A database row can satisfy every type constraint and attach a payment to the wrong account. A spreadsheet can contain every correct number in the wrong cells.
Those failures look similar from the outside because the final artifact is unusable. They are different from the perspective of repair.
If the wrong action was selected, the system may need better retrieval, interpretation, or domain knowledge. If the right action was selected but connected to the wrong predicate, retrieving the action again adds nothing. The missing information is not the component. It is the assignment.
That distinction is likely to matter beyond JSON and workflow graphs wherever roles have checkable identities. A coding task can involve the right files but the wrong dependency update. A research task can contain the right sources but attach one to a claim it does not support. A tool sequence can contain the right calls but route one output into the wrong input. The present studies have not demonstrated those generalizations; they show what to measure if we want to test them.
The cure can become another failure
There is an obvious engineering reaction to binding errors: make every relationship explicit.
An independent study shows why that is also too simple.
Routed Graph Handoff, accepted to EMNLP 2026, compared prose delegations between agents with typed dependency graphs. The graph format improved task success on benchmarks dominated by dependency chains. Then the researchers tried AppWorld, where tasks often require iteration, conditional adaptation, and backtracking.
Graph-only delegation fell from 51.7 percent task success to 37.1 percent—a 14.6-point regression.
The authors traced the difference to rigidity. A dependency edge can prevent an executor from skipping a prerequisite, which is useful when the prerequisite really is fixed. The same edge can become a premature commitment when execution reveals that the plan needs to change. In AppWorld’s iterative and conditional task groups, natural-language handoffs beat graph-only handoffs; graphs helped only the small group of aggregate-query tasks.
Their best system did not declare one representation superior. It routed dependency-chain tasks to graphs and more adaptive tasks to prose. Even the graph itself was not sufficient: the receiving agent needed an accompanying prompt that explained the edge semantics and traversal rules. Passing the same schema without that guidance produced no gain.
This is the boundary the first two studies need.
Explicit structure is valuable when the relationship is both important and stable enough to state correctly. When the relationship is contingent on what happens next, forcing it into a rigid representation can make the agent worse.
The useful design question is therefore not, “Can we encode this relation?” It is, “Do we know this relation well enough, early enough, to make it binding?”
The debugging fork
A single end-to-end score is still necessary when someone has to decide whether an artifact works. It is a poor debugger.
The three studies point to a more informative fork. First ask whether the required pieces were present. Then ask whether the pieces were attached to the right roles.
That separation changes what evidence to collect and what remedy to try. It can reveal a retrieval problem that needs better information, a binding problem that needs better relation handling, or a rigid representation that encoded a relationship before the system had earned the right to treat it as fixed.
The evidence for this distinction is strongest in controlled structured-generation tasks. It does not establish a universal topology bottleneck in coding agents, research agents, or open-ended reasoning. The Microsoft benchmark is manufactured and proprietary; Where vs What deliberately uses synthetic tasks with exact answers; graph handoffs help some task families and hurt others.
But those limitations do not erase the original billing-record failure. They sharpen it.
The model already had the zip code. Retrieving the zip code again would not have repaired the record.
The repair was to put it back where it belonged.
Sources
- Yiwei Zhang, Chengke Wu, Li Wang, and Jianqiang Li, “Where vs What: Decomposing Structural and Content Failures in LLM-Generated Structured Outputs”, arXiv v1, August 26, 2026.
- Anand Iyer, Bhanu Khetharpal, Srinivas Upadhya, and Ramkumar Rajagopal, “Generating Workflow DAGs from Natural Language with Non-Reasoning LLMs”, arXiv v1, August 31, 2026. All four authors are affiliated with Microsoft; the benchmark targets a commercial Microsoft routing-platform dialect and the underlying evaluation data are proprietary.
- Pratyay Banerjee and Ankit Chadha, “Routed Graph Handoff: Adaptive Format Selection for Multi-Agent LLM Delegation”, arXiv v1, August 26, 2026; accepted to EMNLP 2026. The authors are affiliated with Amazon AGI.