prompt-injection · active

Following instructions hidden in data

Models treat instructions found in retrieved documents, tool outputs, or web pages as if they came from the user.

Also called: indirect prompt injection, jailbreak via content

Tags: autonomous-agent, coding-agent, customer-support, rag-qa, security

A capable model treats content it reads through tools as data. It follows only instructions from the user and the system, and it flags text that tries to redirect it rather than acting on it.

Claims

Techniques

Related: Writing secure code and dependencies, Prioritizing safety under conflicting goals
Suggest a change

Capabilities are a way of carving up the subject, and carvings are arguable. Say so if this one is wrong — especially a proposed one, which a pipeline added because several papers used the same framing, not because anyone decided it was right.