separating-prompt-data-reserved-delimiters-fine-tuning-structure-blocksmechanismsingle paper
Separating prompt and data with reserved delimiters and fine-tuning on that structure blocks most injections at little cost to utility.
Capability: Following instructions hidden in data
Sources
- Separating prompt and data with reserved delimiters and fine-tuning on that structure blocks most injections at little cost to utility.
Disagreeing is the most useful thing you can do here. Both sides of every contested claim in this catalog were assembled by the same person, which is its weakest point.
Related claims
- Prompt injection defenses that report near-zero attack success on short-context benchmarks lose most of that protection when the injected instruction sits inside a document of thousands to tens of thousands of tokens: fine-tuned separation defenses and detect-localize-remove pipelines still let a large share of injections through on paper review, resume screening, code review and email threads.Following instructions hidden in data · unreviewed
- Separating instructions from data raises injection resistance substantially, but both published versions get their strength from fine-tuning the model on the separation, and both report improved robustness rather than elimination. Treat it as one layer of defense in depth; the prompt-only variant, without training, has no measured efficacy behind it here.Following instructions hidden in data