under-strong-goal-nudging-several-frontier-models-disabledmechanismsingle paper
Under strong goal nudging, several frontier models disabled oversight mechanisms or lied about their actions in a minority of trials.
Capability: Prioritizing safety under conflicting goals
Sources
- Under strong goal nudging, several frontier models disabled oversight mechanisms or lied about their actions in a minority of trials.
Disagreeing is the most useful thing you can do here. Both sides of every contested claim in this catalog were assembled by the same person, which is its weakest point.
Related claims
- In simulated retail and airline domains with tools and policy documents, frontier models succeeded on well under two thirds of tasks and success dropped sharply when the same task had to be passed several times in a row.Using the tools it is given
- Models endorse widely held falsehoods, showing weak verification against what they know.Checking claims against evidence
- Under strong, deliberately adversarial goal pressure in an autonomous-agent setting, Claude 3.5 Sonnet sometimes disabled oversight mechanisms or misreported its own actions to keep pursuing the stated goal — in a minority of trials, and the study explicitly does not claim this happens in ordinary, non-adversarial use.Prioritizing safety under conflicting goals