Why Tools Can Fail Before an Agent Makes a Choice
Function-calling accuracy and deployment reliability are different measurements: a system must first make tools callable, then recover when they stop behaving as expected.
Haseeb Mohammed Afsar drew 400 servers from one slice of the official Model Context Protocol registry and tried to start them.
Under his protocol, 205 trials ended before an AI model could choose a tool.
The experiment, reported in Afsar’s September 2026 preprint What a Random Draw from the MCP Registry Contains, and What Tool-Use Benchmarks Contain Instead, was deliberately unforgiving. Each sampled package got one launch attempt over MCP’s local standard-input-and-output transport. The probe supplied no credentials, made no repair, and did not retry.
Of the 400 sampled entries, 195 completed initialization. Fifty-three were classified as needing credentials, two packages were unavailable, and 150 fell into the released runner’s residual handshake-failed category.
Those numbers do not mean that half of MCP is broken. The behavioral sample came from active npm-published, stdio-declared servers rather than the whole ecosystem, and a single unrepaired attempt counts transient or configuration-sensitive failures as exclusions. Afsar explicitly describes the 48.8 percent initialization rate as a lower bound on the fraction that could ever start.
But the experiment exposes a boundary that matters for tool-using agents: a great deal can go wrong before model competence becomes the active variable.
The benchmark can start after the environment is clean
A tool-use benchmark has good reasons to remove that mess.
The Berkeley Function Calling Leaderboard, for example, says that it evaluates an LLM’s ability to call functions, or tools, accurately. If the question is which model chooses the right function and arguments, letting one model encounter an outage while another gets a healthy endpoint would contaminate the comparison.
Other benchmarks make the control even more explicit. The authors of StableToolBench built a virtual API server with a cache and simulated responses because live APIs in the benchmark they were extending changed status, expired, failed authentication, or otherwise made repeated evaluation unstable.
That is not a methodological mistake. It is the point of experimental control.
It also changes the question.
A function-calling score can tell us how well a model performs given a usable tool surface. A deployed system must first produce that surface. It may have to install a package, supply configuration or credentials, start a process, negotiate a protocol, survive execution failures, and recover when the environment changes.
A score taken after those conditions have been stabilized is a conditional measurement. Trouble begins only when that conditional result is used as evidence about the reliability of the whole deployed system.
Afsar measured the gate instead
Afsar’s sampling frame contained 7,258 active registry entries that declared npm packaging and stdio transport. A published seed selected 400 without replacement. The probe then recorded an outcome for every draw.
The striking result appears after the startup losses. The 195 servers that initialized advertised 2,766 tools, and the study found zero fatal JSON Schema violations among them.
That does not establish that every surviving interface was semantically good, safe, or useful. It establishes something narrower: the tool definitions that made it through the startup gate passed the study’s hard schema checks.
So the largest problem in this experiment was not malformed function definitions of the kind a model sees once tool calling begins. Much of the attrition happened earlier.
Afsar also ran the same instrument against a hand-curated set of 24 servers. Sixteen initialized, or 66.7 percent, compared with 48.8 percent in the random sample. The direction is suggestive, but the curated set is small and not a probability sample. The 17.9-point difference is not a robust causal estimate of what curation does to the ecosystem.
The stronger observation needs less arithmetic: an unrepaired random draw lost many trials before tool selection, while the servers that survived exposed tool definitions with no fatal schema violations under the study’s checks.
That is enough to show why the denominator matters.
Repair changes what is being measured
A different research strategy makes the same boundary visible from the other side: do not discard hard-to-run servers; repair them.
The 2026 MCPZoo study built a multi-agent system that inferred environments, generated deployment configurations, diagnosed failures, and iterated until a server supported runtime analysis or a retry limit was reached.
Among projects that already had a Dockerfile, README, and configuration artifacts, only 19.6 percent ran successfully without modification in the authors’ setup. With the broader MCPZoo pipeline, global deployment success reached 57.7 percent.
Those percentages are not directly comparable with Afsar’s. The populations, tooling, and success criteria differ. What matters here is the intervention: repair turned some non-runnable projects into runnable ones.
For a security study whose goal is to observe server behavior, that repair is useful research infrastructure. For an end-to-end reliability study, the same repair becomes part of the system being evaluated. A deployment that supplies configuration, restarts a process, or routes around a failed provider may turn a failed first launch into a recoverable event rather than a failed task.
That is the strongest reason not to treat operability as a simple pass-or-fail gate.
Recovery turns the gate into a path
ToolBench-X tests what happens after the clean environment stops being clean.
Its tasks begin with deterministic tools, then introduce five recoverable hazards: specification drift, invocation error, execution failure, output drift, and conflict between sources. Every injected case retains at least one valid recovery path, such as retrying, using a fallback, verifying a result, or cross-checking another source.
Agents that worked well with reliable tools often failed under those hazards. Targeted recovery hints restored many failed tasks. The authors found that failures were driven less by tool-use volume or inference budget than by weak hazard diagnosis and ineffective recovery.
That finding breaks a temptingly simple model of deployment reliability. We cannot always multiply “probability the tool starts” by “probability the model succeeds once it starts” and stop there. A system can retry, reconfigure, substitute another capability, or verify that a suspicious result should not be trusted.
The unit that matters in deployment is often the trajectory: what the system can do after something goes wrong.
That leaves at least three distinct questions inside the phrase tool use:
- Can the system make the required capability callable?
- Given a callable tool, can the model select and invoke it correctly?
- When the environment fails or changes, can the system recover and still produce verified useful work?
One benchmark does not need to answer all three. Collapsing them into one score can make diagnosis worse. The important thing is to keep the claim attached to the denominator that produced it.
Keep both measurements
This distinction changes how an improvement should be interpreted.
A better model might raise function-calling accuracy while accepted task completion barely moves because provisioning, authentication, runtime failure, or recovery has become the bottleneck. The benchmark can be correct and the model genuinely better even when users see little change.
The reverse can happen too. A platform can make more tools operational or recover more failed runs without changing the model at all. A model-only leaderboard would show no improvement even as more work reaches completion.
So a useful evaluation stack should preserve failure ownership rather than hide it in a single tool-use number: how often a required capability becomes callable, how well the model performs once it is callable, how often recovery restores a productive path, and whether the final result passes an independent check.
Afsar’s experiment makes the first boundary unusually concrete. In his 400-server sample, 195 reached an initialized tool surface under the standardized one-shot probe. The other 205 did not.
For those 205 trials, function-calling accuracy was not low.
There was no function call to score.
Sources
- Haseeb Mohammed Afsar, What a Random Draw from the MCP Registry Contains, and What Tool-Use Benchmarks Contain Instead, arXiv, September 10, 2026.
- Haseeb Mohammed Afsar, released
mcp-probestudy code and probe runner, 2026. - Berkeley Function Calling Leaderboard (BFCL) V4, Berkeley Gorilla project.
- Zhicheng Guo et al., StableToolBench: Towards Stable Large-Scale Benchmarking on Tool Learning of Large Language Models, Findings of ACL 2024.
- Pei Chen et al., Rethinking MCP Security: A Large-Scale Study of Runtime MCP Servers and Security Scanner Reliability, arXiv, July 13, 2026.
- Yang Tian, Zhengpeng Shi, and Bo Zhao, Beyond Function Calling: Benchmarking Tool-Using Agents under Tool-Environment Unreliability, arXiv, June 24, 2026.