prompt-injection-defenses-report-near-zero-attack-success-short-context-benchmarks
observationsingle paperpending review

Prompt injection defenses that report near-zero attack success on short-context benchmarks lose most of that protection when the injected instruction sits inside a document of thousands to tens of thousands of tokens: fine-tuned separation defenses and detect-localize-remove pipelines still let a large share of injections through on paper review, resume screening, code review and email threads.

Ingested from a paper but not yet reviewed by a human. It is deliberately inert: it does not move any technique’s standing, does not count toward the backtest, and is excluded anywhere a claim would carry weight. Read the source before relying on it.

Capability: Following instructions hidden in data

Observed on

Single-pass document-centric tasks with full document in one inference call; open 3-8B models plus GPT-4o/GPT-4.1; synthetic and real-world datasets, 100 instances per synthetic su.

Sources

  • Measured: eight LLMs, six heuristic and two optimization attacks, nine detectors and six prevention defenses, four scenarios with synthetic and real-world splits. Baseline is the same defenses on OPI, InjecAgent, AgentDojo, where their reported attack success is near zero. Design shows the effect but the authors state it does not isolate the mechanism (dilution vs position). Detection defenses show extreme false-positive/false-negative tradeoffs. Single group, one benchmark.
Status: pending-reviewLast checked: 2026-09-09Evidence activity: not checked yet
Contest this claim

Disagreeing is the most useful thing you can do here. Both sides of every contested claim in this catalog were assembled by the same person, which is its weakest point.

Related claims

Notes

Drafted from the paper by a model and filed unreviewed. Visible here so it can be read, not because anyone has vouched for it: it does not move any technique's standing and does not count toward the internal scorecard. Drafted confidence: medium. Falsifier as drafted: Evaluating the same defenses on long-document injection with matched attacks and finding attack success stays near the short-context level, or showing the long-context gap disappears once no-attack baseline rates and scoring criteria are matched to the short-context benchmarks. Drafted stance toward separating-prompt-data-reserved-delimiters-fine-tuning-structure-blocks: contests -- A fine-tuned separation defense (MetaSecAlign 8B) that reports near-zero attack success on short-context benchmarks reached attack success of 1.00 on the long-context paper review and resume screening datasets.