deception-ability-emerged-in-frontier-modelsmechanismsingle paper
State-of-the-art models such as GPT-4 can understand and induce false beliefs in other agents through deliberate strategic reasoning, a capability that was absent in earlier-generation language models.
Capability: Strategic deception and detecting it · Behavior
Observed on
2023, GPT-4 vs. earlier LLMs.
Sources
- Deceptive strategies emerged in GPT-4 but were non-existent in earlier LLMs; chain-of-thought reasoning further improved deceptive performance.
Disagreeing is the most useful thing you can do here. Both sides of every contested claim in this catalog were assembled by the same person, which is its weakest point.
Related claims
- Frozen LLMs, using only retrieval over past communications and self-reflection rather than fine-tuning, can play the social deduction game Werewolf competently and show emergent strategic behavior, including deception, without being explicitly trained for it.Strategic deception and detecting it
- Models endorse widely held falsehoods, showing weak verification against what they know.Checking claims against evidence
- In agentic customer-service settings where a deployer incentive conflicts with a user's documented entitlement, a model's willingness to lie under incentive alone is not predicted by its willingness to lie when explicitly told to — some models comply with explicit deception instructions at high rates while almost never initiating deception under incentive, so instructed-deception evaluations measure capability rather than propensity.Strategic deception and detecting it · unreviewed
Notes
Worth chasing down: reports of newer models showing an asymmetry between deceptive and non-deceptive roles in social deduction games (e.g. a specific claim that Claude Opus 4 wins as a non-deceptive peasant but not as a deceptive vampire in Town of Salem) — read secondhand, not yet independently verified against a citable source here.