Skip to content

Fewer Connections, Bigger AI Bills?

In LLM multi-agent systems, communication sparsity can be a misleading cost proxy because removing an edge can also change what downstream models know and therefore what they generate; selective pruning succeeds when it removes low-value information, so executed traffic—not graph density—is the cost that must be measured.

Before the models ran, the cheapest-looking team was easy to spot.

Four agents arranged in a chain need only three direct connections. A fully connected four-agent team needs many more. If communication cost lives in the arrows, the chain should win before anyone spends a token.

Then the researchers ran the teams.

In the new Codebook Agent study, 50 training questions from each benchmark were executed under six fixed graph shapes, producing 300 topology-and-cost records per benchmark. The ledger reversed the diagram. Across those records, edge count was negatively correlated with measured token use, at about r = -0.4. The chain—the sparsest connected fixed topology—was the most token-expensive topology on every math-format benchmark the paper reports. On MATH, it used 2.9 times as many tokens as the fully connected graph.

The result does not show that sparse teams are expensive. Other systems have made sparse teams cheaper.

It shows something more basic: an arrow is not a billable unit.

What an arrow removes

A graph edge looks like a quantity. It is actually a permission.

In the public Codebook Agent implementation, each node gathers the outputs of its graph predecessors before the node runs. Delete an edge and a downstream model may lose intermediate work that it would otherwise have received. The edit changes the model’s input, not merely the number beside the graph.

That matters because a language model does not emit a fixed-size packet. Its response is generated from whatever context it sees. Change that context and the next output can change in content, length, redundancy, and strategy.

The Codebook authors attribute their inverted cost pattern to sparse communication producing longer completions. Their data are consistent with that explanation. But the experiment does not isolate it cleanly enough to make “fewer links cause longer messages” a general law. Topology, available predecessor information, prompt construction, and generated output all change together.

A decisive experiment would remove a direct edge while restoring the missing information through some other channel. Codebook Agent does not do that. So the strongest verified claim is one step earlier: changing the graph changes the information conditions under which the remaining model calls execute.

Once that is true, edge count no longer has a fixed exchange rate with tokens.

Two sparse graphs can mean opposite things

The strongest evidence against an anti-sparsity story comes from systems built to prune communication deliberately.

AgentPrune, published at ICLR 2025, identifies communication it considers redundant and reports token reductions of 28.1 to 72.8 percent across six benchmarks. AgentDropout, published at ACL 2025, dynamically removes agents and connections and reports average reductions of 21.6 percent in prompt tokens and 18.4 percent in completion tokens.

Those results are not awkward exceptions to Codebook Agent. They clarify it.

A chain is sparse by shape. Selective pruning is sparse by judgment.

Imagine two graphs with the same number of edges. One removes repeated information while preserving the links that carry unique intermediate results. The other removes the unique results and keeps the repetition. Their density is identical. Their informational value is not.

Edge count cannot distinguish them.

This is the flaw in using density as a general cost proxy. The number of surviving paths says nothing about whether the removed paths were redundant, indispensable, or replaceable somewhere else.

The payload can move

There is a second reason token traffic cannot be read from topology alone: the payload is trainable.

Optima, published in Findings of ACL 2025, trains multi-agent communication with a reward that includes task performance, token efficiency, and readability. On some information-intensive tasks, its authors report much higher performance while using less than a tenth of the tokens of comparison systems in their tested setup.

That does not prove the Codebook chain agents were reconstructing missing context. It establishes a broader systems fact: how much text a multi-agent system produces is partly a property of its communication policy, not merely its node and edge count.

This is why a structural variable can be useful for design without being a meter for execution. Graph density tells an orchestrator something real about who may communicate. It does not, by itself, tell the orchestrator how much prompt material will be repeated, how verbose downstream calls will become, or whether an agent will have to rediscover information that disappeared with a link.

The distinction becomes even more important in software work. Coding agents can exchange state through direct messages, but also through files, commits, tests, issue state, plans, and other durable artifacts. Removing a conversational edge may remove information, or merely change its route. The Codebook experiment does not test those software-work channels, so claims about Git churn, retries, or human review remain hypotheses.

But it tells us exactly what to ask: what disappeared with the edge, and where did the work go afterward?

A young result with a useful boundary

Codebook Agent is a September 2, 2026 arXiv v1, not an independently replicated result. Its main experiments use small benchmark teams; six of seven detailed design settings use homogeneous agent profiles. The main backbone is gpt-4o-mini, with Qwen-3-8B used as a second-backbone check. The paper reports a single evaluation run per configuration, and its public repository does not include the exact sampled benchmark subsets needed to reproduce every absolute number from the published setup.

Those limits make the 2.9-fold MATH result a research finding, not a universal coefficient for sparse communication.

They do not erase the measurement failure it reveals. In the authors’ own execution records, a structural variable chosen to stand in for cost pointed in the opposite direction from the cost actually measured.

The right response is not to replace “fewer edges are cheaper” with “fewer edges are more expensive.” It is to stop asking graph density to answer a question it does not contain enough information to answer.

Measure the execution

An orchestration system can still use sparsity as a design constraint. It can still prefer simple graphs, discourage redundant fan-out, or search for fewer handoffs. But if the objective is cost, the cost needs its own instrumentation.

Measure prompt tokens and completion tokens separately. Measure latency if latency matters. For coding agents, include retries, tool calls, repeated repository reads, integration work, and reviewer effort when those are part of the bill. Compare equal-density graphs that remove different information, because that is where a raw edge count is most likely to conceal the mechanism.

The Codebook chain looked cheap before execution because three arrows are easy to count.

But in its executor, deleting an arrow can also delete an input to the next model call. Before a topology optimizer rewards the missing edge, it needs the ledger for the execution that edge changed.

Sources