training-models-rank-system-user-tool-instructions-by
mechanismsingle paper

Training models to rank system, user, and tool instructions by privilege improves robustness to injections in tool outputs.

Capability: Following instructions hidden in data

Sources

Status: activeLast checked: 2026-09-03Evidence activityHow much the field cites the sources under this claimvery heavily cited in the last 12 months285 in 12mo · 486 total — The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions
Contest this claim

Disagreeing is the most useful thing you can do here. Both sides of every contested claim in this catalog were assembled by the same person, which is its weakest point.

Related claims