deception-ability-emerged-in-frontier-models
mechanismsingle paper

State-of-the-art models such as GPT-4 can understand and induce false beliefs in other agents through deliberate strategic reasoning, a capability that was absent in earlier-generation language models.

Capability: Strategic deception and detecting it · Behavior

Observed on

2023, GPT-4 vs. earlier LLMs.

Sources

Status: activeLast checked: 2026-09-04Evidence activityHow much the field cites the sources under this claimheavily cited in the last 12 months75 in 12mo · 191 total — Deception Abilities Emerged in Large Language Models
Contest this claim

Disagreeing is the most useful thing you can do here. Both sides of every contested claim in this catalog were assembled by the same person, which is its weakest point.

Related claims

Notes

Worth chasing down: reports of newer models showing an asymmetry between deceptive and non-deceptive roles in social deduction games (e.g. a specific claim that Claude Opus 4 wins as a non-deceptive peasant but not as a deceptive vampire in Town of Salem) — read secondhand, not yet independently verified against a citable source here.