removing-overconstraint-from-system-prompts-stopped-costing-accuracy
observationsingle paperpending review

Guardrail instructions written to stop older models doing the wrong thing — blanket prohibitions, repeated warnings, defensive defaults — became dead weight in the Claude 5 generation: Anthropic removed over 80% of Claude Code's system prompt for Opus 5 and Fable 5 with no measurable loss on its own coding evaluations, on the reading that the constraints now conflict with each other and with user intent more often than they prevent harm.

Ingested from a paper but not yet reviewed by a human. It is deliberately inert: it does not move any technique’s standing, does not count toward the backtest, and is excluded anywhere a claim would carry weight. Read the source before relying on it.

Capability: Keeping its own context clean · Context and memory, Coding agent

Observed on

2026, Claude Opus 5 and Fable 5 in Claude Code.

Sources

  • Reports the 80% reduction and the reasoning: transcripts showed system prompt, skills and user request issuing conflicting instructions in a single request, which the model then had to spend effort reconciling. The measurement is real but unverifiable from outside — the evaluations are internal and unpublished, so there is no baseline, ablation or independent check available.
Status: pending-reviewLast checked: 2026-09-08Evidence activity: not checked yet
Contest this claim

Disagreeing is the most useful thing you can do here. Both sides of every contested claim in this catalog were assembled by the same person, which is its weakest point.

Related claims

Notes

Filed as an observation about one model family at one moment, not a mechanism. The durable question underneath — whether instruction volume trades against instruction-following, and where the turn happens — is not answered here. Capped at single-paper because the source reports an actual measurement. Vendor guidance that only asserts, without a number, cannot reach even that. And "no measurable loss" is doing quiet work: it means no loss their evaluations could see, which is a weaker statement than no loss.