models-complied-insecure-completions-large-fraction-time-moremechanismsingle paper
Models complied with insecure completions a large fraction of the time, and more capable models were more likely to suggest insecure code.
Capability: Writing secure code and dependencies
Sources
- Models complied with insecure completions a large fraction of the time, and more capable models were more likely to suggest insecure code.
Disagreeing is the most useful thing you can do here. Both sides of every contested claim in this catalog were assembled by the same person, which is its weakest point.
Related claims
- Users with an AI assistant wrote less secure code and were more confident it was secure.Writing secure code and dependencies
- On code prompts that are unsatisfiable by design, open-weight code and reasoning models refuse based on how suspicious the request looks rather than how deeply impossible it is: they refuse famous theoretical impossibilities far more often than requests for plausible-sounding nonexistent packages, and the stronger models in the set close the gap on theory while barely improving on fabricated ecosystem entities.Stating false facts confidently · unreviewed
- Roughly forty percent of Copilot completions in security-relevant scenarios were vulnerable.Writing secure code and dependencies
- On verifiable instructions such as length and format constraints, strong models still fail a meaningful share, and failures grow when several constraints apply at once.Following an unfamiliar procedure
- Guardrail instructions written to stop older models doing the wrong thing — blanket prohibitions, repeated warnings, defensive defaults — became dead weight in the Claude 5 generation: Anthropic removed over 80% of Claude Code's system prompt for Opus 5 and Fable 5 with no measurable loss on its own coding evaluations, on the reading that the constraints now conflict with each other and with user intent more often than they prevent harm.Keeping its own context clean · unreviewed