tool-use-fine-tuning-closes-a-gap-rather-than-extending-a-frontier
observationsingle paper

The large reported gains from fine-tuning on tool-call trajectories come from small open models starting far behind — distilling a few hundred trajectories from a stronger model moves them a long way. Whether the same procedure helps a model that is already good at tool use is a separate question this evidence does not answer.

Capability: Using the tools it is given · Agentic, Coding agent

Observed on

2023, Llama2-7B fine-tuned on GPT-4 trajectories.

Sources

  • Fine-tuning Llama2-7B on 500 GPT-4-generated agent trajectories gave a 77% HotpotQA improvement, and mixing trajectories from multiple tasks and prompting methods improved agents further.
Status: activeLast checked: 2026-09-04Evidence activityHow much the field cites the sources under this claimheavily cited in the last 12 months101 in 12mo · 252 total — FireAct: Toward Language Agent Fine-tuning
Contest this claim

Disagreeing is the most useful thing you can do here. Both sides of every contested claim in this catalog were assembled by the same person, which is its weakest point.

Related claims

Notes

Deliberately stated as a scope limit rather than as "gains are largest for weak models," which would be a comparative the source does not establish — it reports a large gain for a weak model, which is not the same as measuring the gain shrinking as the baseline rises. If a paper ablates starting strength directly, this claim should be tightened or replaced.