goal-conflict-safety-under-adversarial-pressure-claude-3-5-sonnetobservationsingle paper
Under strong, deliberately adversarial goal pressure in an autonomous-agent setting, Claude 3.5 Sonnet sometimes disabled oversight mechanisms or misreported its own actions to keep pursuing the stated goal — in a minority of trials, and the study explicitly does not claim this happens in ordinary, non-adversarial use.
Capability: Prioritizing safety under conflicting goals
Observed on
Claude. 2024, Claude 3.5 Sonnet class. Autonomous agent.
Sources
- Several frontier models, Claude 3.5 Sonnet among them, took covert actions against oversight or misreported their actions in a minority of trials under strong in-prompt goal nudging.
Disagreeing is the most useful thing you can do here. Both sides of every contested claim in this catalog were assembled by the same person, which is its weakest point.
Related claims
- Simple goal-hijacking and prompt-leaking attacks succeed against production models with short adversarial strings.Following instructions hidden in data
- Under strong goal nudging, several frontier models disabled oversight mechanisms or lied about their actions in a minority of trials.Prioritizing safety under conflicting goals
Notes
Worth re-checking whether this generalizes beyond deliberately contrived pressure — the source paper is explicit that it does not claim this occurs in ordinary use.