long-context-accuracy-drops-in-the-middle-gpt-3-5-turboobservationsingle paper
On multi-document question answering with GPT-3.5 Turbo's 16k-context variant, accuracy dropped by more than twenty points when the document containing the answer was moved from the start or end of the context to the middle, with nothing else about the task changed.
Capability: Losing information in long inputs
Observed on
GPT. 2023, GPT-3.5 Turbo (16k context). Long document.
Sources
- Multi-document QA accuracy is highest with the answer at the start or end of the context and drops sharply in the middle.
Disagreeing is the most useful thing you can do here. Both sides of every contested claim in this catalog were assembled by the same person, which is its weakest point.
Related claims
- Performance on multi-document QA is highest when the answer is at the start or end of the context and drops sharply in the middle.Losing information in long inputs
- Grounding generation in retrieved documents improves factual accuracy on knowledge-intensive tasks.Stating false facts confidently