arxiv-2309-04658 · paper

Exploring Large Language Models for Communication Games: An Empirical Study on Werewolf

Yuzhuang Xu, Shuo Wang, Peng Li, Fuwen Luo, Xiaolong Wang, Weidong Liu, Yang Liu

Created: 2023 · Ingested: 2026-09-04

This source: Exploring Large Language Models for Communication Games: An Empirical …How much the field cites it — heavily cited in the last 12 months59 in the last 12 months · 312 totalpublished 2023checked 2026-09-04click for Semantic Scholar59 citations in the last 12 months · 312 total · checked 2026-09-04

https://arxiv.org/abs/2309.04658(opens in a new tab)

Frozen LLMs, using only retrieval over past communications and self-reflection rather than fine-tuning, can play the social deduction game Werewolf competently and show emergent strategic behavior, including deception, without being explicitly trained for it.

In brief

Frozen LLMs can play a 7-player Werewolf game and pick up strategy from past games through retrieval and reflection alone, but the gains do not scale with how much experience you give them. Agents built on gpt-3.5-turbo-0301, one per role (2 werewolves, 2 villagers, witch, guard, seer), get a prompt containing the most recent 15 messages, rule-matched informative messages, a reflection built by answering 5 predefined plus 2 freely asked questions over retrieved memory, a suggestion mined from an experience pool, and a chain-of-thought trigger.

Experience pools were built from 10, 20, 30, and 40 rounds and given only to the village side, with werewolves held fixed as a reference; each condition ran 50 rounds. Pools of 10 and 20 rounds raised villager win rate and game duration; 30 rounds lengthened games without a clear win-rate change; 40 rounds gave shorter games. The paper reports these as chart trends and gives no win-rate numbers.

Evidence is thin. Ablations are mostly qualitative, with 1 human-annotated comparison of 50 sampled responses judged reasonable or not. The werewolf baseline is not actually constant: werewolf camouflage behavior also shifted, so the control is compromised, as the authors concede. 1 model, 1 game, 1 configuration.

Emergent trust, confrontation, camouflage, and leadership are documented by example and by a 7x7 trust table, not measured against a baseline; they persisted when role names were replaced with unrelated words.

Treat this as an existence proof for prompt-only social play, not a scaling result.

Written from the abstract by claude-opus-5 on 2026-09-05, with every figure checked against it. Not a substitute for the paper.

Referenced by

Claims in this catalog that draw on this source, and whether as support or counterpoint.