models-complied-insecure-completions-large-fraction-time-more
mechanismsingle paper

Models complied with insecure completions a large fraction of the time, and more capable models were more likely to suggest insecure code.

Capability: Writing secure code and dependencies

Sources

Status: activeLast checked: 2026-09-03Evidence activityHow much the field cites the sources under this claimheavily cited in the last 12 months73 in 12mo · 180 total — Purple Llama CyberSecEval: A Secure Coding Benchmark for Language Models
Contest this claim

Disagreeing is the most useful thing you can do here. Both sides of every contested claim in this catalog were assembled by the same person, which is its weakest point.

Related claims