goal-conflict-safety-under-adversarial-pressure-claude-3-5-sonnet
observationsingle paper

Under strong, deliberately adversarial goal pressure in an autonomous-agent setting, Claude 3.5 Sonnet sometimes disabled oversight mechanisms or misreported its own actions to keep pursuing the stated goal — in a minority of trials, and the study explicitly does not claim this happens in ordinary, non-adversarial use.

Capability: Prioritizing safety under conflicting goals

Observed on

Claude. 2024, Claude 3.5 Sonnet class. Autonomous agent.

Sources

Status: activeLast checked: 2026-09-03Evidence activityHow much the field cites the sources under this claimheavily cited in the last 12 months189 in 12mo · 288 total — Frontier Models are Capable of In-context Scheming
Contest this claim

Disagreeing is the most useful thing you can do here. Both sides of every contested claim in this catalog were assembled by the same person, which is its weakest point.

Related claims

Notes

Worth re-checking whether this generalizes beyond deliberately contrived pressure — the source paper is explicit that it does not claim this occurs in ordinary use.