models-endorse-widely-held-falsehoods-showing-weak-verificationmechanismsingle paper
Models endorse widely held falsehoods, showing weak verification against what they know.
Capability: Checking claims against evidence
Sources
- Models endorse widely held falsehoods, showing weak verification against what they know.
Disagreeing is the most useful thing you can do here. Both sides of every contested claim in this catalog were assembled by the same person, which is its weakest point.
Related claims
- Models reproduce common human misconceptions, and larger models were not more truthful on this benchmark.Stating false facts confidently
- On questions whose premises are false or internally contradictory, prompting mid-size instruct models (Qwen2.5-7B, LLaMA-3.1-8B, Gemma3-12B) to reason step by step can lower accuracy below plain answering, because the reasoning chain elaborates from the flawed premise instead of challenging it.Stating false facts confidently · unreviewed
- On verifiable instructions such as length and format constraints, strong models still fail a meaningful share, and failures grow when several constraints apply at once.Following an unfamiliar procedure
- Models readily adopt a single counter-memory passage when it is coherent, and when sources conflict they follow the majority and show confirmation bias toward their own beliefs.Checking claims against evidence
- State-of-the-art models such as GPT-4 can understand and induce false beliefs in other agents through deliberate strategic reasoning, a capability that was absent in earlier-generation language models.Strategic deception and detecting it