across-eight-interactive-environments-models-fail-mostly-long-horizonmechanismsingle paper
Across eight interactive environments, models fail mostly on long-horizon tool interaction, with open models far behind commercial ones.
Capability: Using the tools it is given
Sources
- Across eight interactive environments, models fail mostly on long-horizon tool interaction, with open models far behind commercial ones.
Disagreeing is the most useful thing you can do here. Both sides of every contested claim in this catalog were assembled by the same person, which is its weakest point.
Related claims
- For locally deployed open-weight models driving a stateful, dependency-ordered MCP tool server, cutting tool descriptions from full specifications (purpose, parameter semantics, constraints, failure conditions) down to one sentence each raises the fraction of calls the server rejects for every model tested, while its effect on task coverage is less consistent.Using the tools it is given · unreviewed
- For long-horizon tool-calling agents (tasks needing five or more calls), embedding state-transition cues — preconditions, invariants, completion states — into the descriptions of the tools on the intended chain is what actually steers the agent's trajectory; runtime corrective text appended to tool results only patches residual drift, and the plausible user prompt alone (persona, deadlines, format constraints) does not establish the trajectory at all.Following instructions hidden in data · unreviewed
- Fine-tuning on a large corpus of real API-call traces markedly improves multi-tool task completion in open models.Using the tools it is given
- Across many agent environments, most tested models scored below sixty percent on safety, with failures concentrated in unsafe tool actions.Prioritizing safety under conflicting goals
- Models that predict next steps well can still hold an incoherent implicit world model, which fails when the task deviates from familiar traces.Tracking state through a long task