accuracy-fact-scales-how-many-pretraining-documents-mentionmechanismsingle paper
Accuracy on a fact scales with how many pretraining documents mention it, so long-tail facts are systematically unreliable.
Capability: Stating false facts confidently
Sources
- Accuracy on a fact scales with how many pretraining documents mention it, so long-tail facts are systematically unreliable.
Disagreeing is the most useful thing you can do here. Both sides of every contested claim in this catalog were assembled by the same person, which is its weakest point.
Related claims
- Models reproduce common human misconceptions, and larger models were not more truthful on this benchmark.Stating false facts confidently
- Most models claiming long contexts fail well before their advertised length on synthetic retrieval, tracing and aggregation tasks.Losing information in long inputs
- Where a passage sits in a long input changes how much the model uses it — accuracy is highest when the needed information is at the very start or the very end and lowest when it is in the middle — so ordering retrieved passages to put the most relevant ones first is a real lever on accuracy.Losing information in long inputs · contested
- Models endorse widely held falsehoods, showing weak verification against what they know.Checking claims against evidence
- Training the model to decide when to retrieve and to critique whether passages support its output improves factuality and citation accuracy.Checking claims against evidence