frozen-llms-play-werewolf-competently-via-reflection
observationsingle paper

Frozen LLMs, using only retrieval over past communications and self-reflection rather than fine-tuning, can play the social deduction game Werewolf competently and show emergent strategic behavior, including deception, without being explicitly trained for it.

Capability: Strategic deception and detecting it · Behavior, Agentic

Observed on

2023, GPT-3.5/GPT-4 class. Autonomous agent.

Sources

Status: activeLast checked: 2026-09-04Evidence activityHow much the field cites the sources under this claimheavily cited in the last 12 months59 in 12mo · 312 total — Exploring Large Language Models for Communication Games: An Empirical Study on Werewolf
Contest this claim

Disagreeing is the most useful thing you can do here. Both sides of every contested claim in this catalog were assembled by the same person, which is its weakest point.

Related claims