training-against-explicit-set-principles-produces-models-more
mechanismsingle paper

Training against an explicit set of principles produces models that are more harmless without becoming evasive.

Capability: Prioritizing safety under conflicting goals

Sources

Status: activeLast checked: 2026-09-03Evidence activityHow much the field cites the sources under this claimvery heavily cited in the last 12 months1000+ in 12mo · 3617 total — Constitutional AI: Harmlessness from AI Feedback
Contest this claim

Disagreeing is the most useful thing you can do here. Both sides of every contested claim in this catalog were assembled by the same person, which is its weakest point.

Related claims