most-models-claiming-long-contexts-fail-well-beforemechanismsingle paper
Most models claiming long contexts fail well before their advertised length on synthetic retrieval, tracing and aggregation tasks.
Capability: Losing information in long inputs
Sources
- Most models claiming long contexts fail well before their advertised length on synthetic retrieval, tracing and aggregation tasks.
Disagreeing is the most useful thing you can do here. Both sides of every contested claim in this catalog were assembled by the same person, which is its weakest point.
Related claims
- Accuracy on a fact scales with how many pretraining documents mention it, so long-tail facts are systematically unreliable.Stating false facts confidently
- Fine-tuning on a large corpus of real API-call traces markedly improves multi-tool task completion in open models.Using the tools it is given
- A tiered memory system managed by the model itself sustains recall over conversations far longer than the context window.Remembering across sessions
- Only models with substantial code pretraining track entity state through a sequence of operations, and all degrade as the sequence lengthens.Tracking state through a long task
- When retrieval for a summarization task pulls definitions by surface term match from a general encyclopedia, ambiguous terms fetch the wrong sense and the model incorporates that unsupported background into its output, lowering claim-level factuality below the no-retrieval baseline even when the prompt tells it to ignore unrelated retrieved content.Stating false facts confidently · unreviewed