locally-deployed-open-weight-models-driving-stateful-dependency-ordered-mcp-tool
observationsingle paperpending review

For locally deployed open-weight models driving a stateful, dependency-ordered MCP tool server, cutting tool descriptions from full specifications (purpose, parameter semantics, constraints, failure conditions) down to one sentence each raises the fraction of calls the server rejects for every model tested, while its effect on task coverage is less consistent.

Ingested from a paper but not yet reviewed by a human. It is deliberately inert: it does not move any technique’s standing, does not count toward the backtest, and is excluded anywhere a claim would carry weight. Read the source before relying on it.

Capability: Using the tools it is given

Observed on

Seven 4-bit quantised open models (8B-31B) run through Ollama on an expert-informed hardware-design benchmark of 14 tools; single-agent ReAct loop with other configuration held fix.

Sources

  • Measured as tool failure rate under matched configurations; the paper reports TFR rises for every model and roughly doubles for most, and comprehensive descriptions win in 35 of 42 model-suite best configurations. No absolute per-model numbers given for the ablation, and only one tool set of 14 tools on a proprietary replica server.
Status: pending-reviewLast checked: 2026-09-09Evidence activity: not checked yet
Contest this claim

Disagreeing is the most useful thing you can do here. Both sides of every contested claim in this catalog were assembled by the same person, which is its weakest point.

Related claims

Notes

Drafted from the paper by a model and filed unreviewed. Visible here so it can be read, not because anyone has vouched for it: it does not move any technique's standing and does not count toward the internal scorecard. Drafted confidence: medium. Falsifier as drafted: An ablation on comparable stateful tool benchmarks where minimal one-sentence descriptions leave the tool failure rate unchanged or lower than comprehensive descriptions across models. Drafted stance toward shows-llms-hallucinate-api-names-arguments-when-calling: supports -- Both find that giving the model detailed tool documentation reduces malformed or invalid calls.