strategic-deception · active

Strategic deception and detecting it

Whether a model can deliberately induce false beliefs in others when a task calls for it, and detect deception from others.

Tags: behavior

Cleanly testable in social deduction games (Werewolf, Mafia, Among Us), where deception is the explicit task rather than an unwanted side effect. This is a capability, not the same question as goal-conflict-safety: that capability asks whether a model deceives its own principal against instructions (a safety property); this one asks how skilled the model is at deception when the task legitimately calls for it. Knowing a model is bad at deception is reassuring context for the safety question — the two are worth reading together.

Claims

Techniques

None yet.

Related: Prioritizing safety under conflicting goals
Suggest a change

Capabilities are a way of carving up the subject, and carvings are arguable. Say so if this one is wrong — especially a proposed one, which a pipeline added because several papers used the same framing, not because anyone decided it was right.