arxiv-2109-07958 · paperTruthfulQA: Measuring How Models Mimic Human Falsehoods
Stephanie Lin, Jacob Hilton, Owain Evans
Created: 2021 · Ingested: 2026-09-02
This source: TruthfulQA: Measuring How Models Mimic Human FalsehoodsHow much the field cites it — very heavily cited in the last 12 months1000+ in the last 12 months · 3851 totalpublished 2021checked 2026-09-04count capped at 1000 by the fetch1000+ citations in the last 12 months · 3851 total · checked 2026-09-04
https://arxiv.org/abs/2109.07958(opens in a new tab)In brief
Bigger language models are less truthful on questions designed to elicit common human misconceptions, reversing the usual scaling trend. The reason offered is that false answers with high likelihood on web text -- "imitative falsehoods" -- get learned better as models improve.
TruthfulQA is 817 hand-written questions across 38 categories (health, law, finance, politics), scored zero-shot on full-sentence generation with human judges, plus a multiple-choice variant. Tested families: GPT-3, GPT-Neo/J, GPT-2, and UnifiedQA (T5), across sizes, with several prompts for GPT-3-175B.
The best model, GPT-3-175B with a helpful prompt, was truthful on 58% of questions against 94% for a human participant, and gave false-but-informative answers 42% of the time versus 6% for the human. The largest GPT-Neo/J was 17% less truthful than a model 60x smaller; on multiple choice, GPT-Neo/J 6B was 12% less truthful than the 125M version and no model beat random guessing.
The imitative explanation is supported rather than assumed: matched control questions edited by 1-3 words improve with scale, paraphrases reproduce the same falsehoods, and the trend transfers to GPT-Neo/J, which was not used for adversarial filtering. Questions and reference answers were written by the authors; external validators disagreed on 6-7%.
437 of the questions were adversarially filtered against GPT-3-175B, so absolute scores understate general truthfulness. Nothing covers long-form generation or dialogue.
Treat inverse scaling here as evidence about imitation objectives, not about model capacity.
Written from the abstract by claude-opus-5 on 2026-09-05, with every figure checked against it. Not a substitute for the paper.
Referenced by
Claims in this catalog that draw on this source, and whether as support or counterpoint.