jailbreak-robustness · proposedRefusing harmful requests under obfuscation
Proposed. Several papers in the review queue converged on this framing, so the pipeline added it. Nobody has decided it is the right way to carve up the subject — it may be two topics, or a duplicate of another, or not a topic at all. Saying so is useful.
Covers maintaining refusal when harmful intent is wrapped in roleplay or encodings, harm that emerges only from combined segments, and unsafe content inside reasoning traces; excludes instructions injected via retrieved data.
Tags: general
Proposed from the ingestion pipeline rather than chosen by hand. 3 papers in the review queue independently pointed at this same competence, arriving under 3 different names (jailbreak-robustness, reasoning-trace-safety, compositional-harm-detection), which is the signal that it is a real recurring topic and not one author's framing. Complements prompt-injection and goal-conflict-safety without duplicating them, with three supporting papers.
What counts as this capability
Scope boundary used when deciding whether a paper is really about this capability, rather than merely mentioning it.
Covers maintaining refusal when harmful intent is wrapped in roleplay or encodings, harm that emerges only from combined segments, and unsafe content inside reasoning traces; excludes instructions injected via retrieved data.
Claims
No claims filed yet.
Techniques
None yet.
Capabilities are a way of carving up the subject, and carvings are arguable. Say so if this one is wrong — especially a proposed one, which a pipeline added because several papers used the same framing, not because anyone decided it was right.