training · constitutional-training

Train against explicit principles

Use a written set of principles to generate critiques and preferences, then train on them.

The model critiques and revises its own outputs against a list of principles, and a preference model trained on those judgments guides reinforcement learning. Produces models that refuse harmful actions while explaining why rather than going evasive.

Requires: A vendor-side training method: application developers cannot apply it directly, so it explains differences between vendors rather than being something you deploy.

Addresses: Prioritizing safety under conflicting goals

Contexts: Chat assistant, Autonomous agent

Does it work?

nothing measured0 supporting · 0 contesting sources

Efficacy claims — what this technique actually moves, under which conditions, and whether that has been contested.

No efficacy claim filed yet. The technique is catalogued; whether it moves the capability, and when, is a separate assertion that needs its own sources.

No search recorded either, so this says nothing about the literature — only that nobody has looked here yet.

Code

No repository linked yet. Contribute one.

Sources