long-context-accuracy-drops-in-the-middle-gpt-3-5-turbo
observationsingle paper

On multi-document question answering with GPT-3.5 Turbo's 16k-context variant, accuracy dropped by more than twenty points when the document containing the answer was moved from the start or end of the context to the middle, with nothing else about the task changed.

Capability: Losing information in long inputs

Observed on

GPT. 2023, GPT-3.5 Turbo (16k context). Long document.

Sources

Status: activeLast checked: 2026-09-03Evidence activityHow much the field cites the sources under this claimvery heavily cited in the last 12 months1000+ in 12mo · 4887 total — Lost in the Middle: How Language Models Use Long Contexts
Contest this claim

Disagreeing is the most useful thing you can do here. Both sides of every contested claim in this catalog were assembled by the same person, which is its weakest point.

Related claims