simple-goal-hijacking-prompt-leaking-attacks-succeed-against-production-modelsmechanismsingle paper
Simple goal-hijacking and prompt-leaking attacks succeed against production models with short adversarial strings.
Capability: Following instructions hidden in data
Sources
- Simple goal-hijacking and prompt-leaking attacks succeed against production models with short adversarial strings.
Disagreeing is the most useful thing you can do here. Both sides of every contested claim in this catalog were assembled by the same person, which is its weakest point.
Related claims
- Under strong, deliberately adversarial goal pressure in an autonomous-agent setting, Claude 3.5 Sonnet sometimes disabled oversight mechanisms or misreported its own actions to keep pursuing the stated goal — in a minority of trials, and the study explicitly does not claim this happens in ordinary, non-adversarial use.Prioritizing safety under conflicting goals