putting-relevant-passages-first-mitigates-position-effects
mechanismsingle papercontested

Where a passage sits in a long input changes how much the model uses it — accuracy is highest when the needed information is at the very start or the very end and lowest when it is in the middle — so ordering retrieved passages to put the most relevant ones first is a real lever on accuracy.

Capability: Losing information in long inputs · Context and memory, Retrieval-augmented QA

Observed on

2023, GPT-3.5 / Claude / open long-context models.

Sources

  • Measures the U-shaped curve directly: multi-document QA and key-value retrieval accuracy drops sharply when the relevant document sits in the middle of the context, and is highest at the beginning or the end.
  • Contests the remedy rather than the effect. When several passages genuinely conflict, whichever is placed first dominates the output, and randomly permuting the evidence does not move that dominance — the skew is already present in the residual-stream representation of the combined prompt, so reordering cannot be the fix.
  • A third position, from the builder. Anthropic reports that earlier models "could sometimes need repeated instructions or be more likely to listen to instructions at the end of their context window than at the start", and that for the Claude 5 generation they deleted those repeats with no measurable loss. If the positional effect has weakened in newer models, the technique is a remedy for a property that is being trained away rather than a durable one. Unverifiable from outside: the evaluations are not published.

Suspected axis of disagreement(a guess, not verified)

Two axes now, because there are three positions. The first is what the other passages are. The supporting result places one relevant passage among distractors, where moving it out of the middle recovers accuracy that position was costing. The 2026 conflict result places several mutually incompatible passages of equal legitimacy, where position is not degrading access to a known answer but deciding which answer wins — and reordering only moves which passage occupies the winning slot. If that is right, both hold, and ordering is a remedy for distraction rather than for conflict. The second is model generation, and it cuts across the first. Anthropic says the end-of-context bias it used to prompt around is largely gone in the Claude 5 generation. So the effect may be real, real still under conflict, and being trained away for ordinary retrieval — three claims that can all hold at once about different models at different times. This is exactly what the mechanism-versus-observation split is for, and it is an argument that this claim is closer to observation than it is currently filed. I have not tested any of the three setups against each other.

Status: activeLast checked: 2026-09-07Evidence activityHow much the field cites the sources under this claimvery heavily cited in the last 12 months1000+ in 12mo · 4887 total — Lost in the Middle: How Language Models Use Long Contexts2 sources not yet checked, so not counted
Contest this claim

Disagreeing is the most useful thing you can do here. Both sides of every contested claim in this catalog were assembled by the same person, which is its weakest point.

Related claims

Notes

Filed deliberately as an incumbent position rather than as a held belief. The catalog previously recorded only claims it endorsed, which left contesting evidence with nothing to attach to — the reason the backtest grades known reversals as missed even when the reversing paper is on file. A claim someone reasonably holds, with the evidence against it attached, is more useful than its absence.