constitutional-trainingTrain against explicit principles
Use a written set of principles to generate critiques and preferences, then train on them.
The model critiques and revises its own outputs against a list of principles, and a preference model trained on those judgments guides reinforcement learning. Produces models that refuse harmful actions while explaining why rather than going evasive.
Requires: A vendor-side training method: application developers cannot apply it directly, so it explains differences between vendors rather than being something you deploy.
Contexts: Chat assistant, Autonomous agent
Does it work?
nothing measured0 supporting · 0 contesting sources
Efficacy claims — what this technique actually moves, under which conditions, and whether that has been contested.
No efficacy claim filed yet. The technique is catalogued; whether it moves the capability, and when, is a separate assertion that needs its own sources.
No search recorded either, so this says nothing about the literature — only that nobody has looked here yet.
Code
No repository linked yet. Contribute one.