simple-goal-hijacking-prompt-leaking-attacks-succeed-against-production-models
mechanismsingle paper

Simple goal-hijacking and prompt-leaking attacks succeed against production models with short adversarial strings.

Capability: Following instructions hidden in data

Sources

Status: activeLast checked: 2026-09-03Evidence activityHow much the field cites the sources under this claimvery heavily cited in the last 12 months475 in 12mo · 1055 total — Ignore Previous Prompt: Attack Techniques For Language Models
Contest this claim

Disagreeing is the most useful thing you can do here. Both sides of every contested claim in this catalog were assembled by the same person, which is its weakest point.

Related claims