training-against-explicit-set-principles-produces-models-moremechanismsingle paper
Training against an explicit set of principles produces models that are more harmless without becoming evasive.
Capability: Prioritizing safety under conflicting goals
Sources
- Training against an explicit set of principles produces models that are more harmless without becoming evasive.
Disagreeing is the most useful thing you can do here. Both sides of every contested claim in this catalog were assembled by the same person, which is its weakest point.
Related claims
- Training on simple gameable environments generalizes, rarely, to tampering with the model's own reward mechanism.Prioritizing safety under conflicting goals
- Models endorse widely held falsehoods, showing weak verification against what they know.Checking claims against evidence
- Separating instructions from data raises injection resistance substantially, but both published versions get their strength from fine-tuning the model on the separation, and both report improved robustness rather than elimination. Treat it as one layer of defense in depth; the prompt-only variant, without training, has no measured efficacy behind it here.Following instructions hidden in data
- In open-weight instruction-tuned models, interventions that suppress caving to user pushback (DPO, SFT on chosen responses, or activation steering) also tend to reduce the model's rate of correcting a wrong answer when genuine supporting evidence is supplied, because the two answer-flip behaviors run on overlapping MLP neurons and attention heads with positively aligned steering directions.Telling the user what they want to hear · unreviewed
- Models complied with insecure completions a large fraction of the time, and more capable models were more likely to suggest insecure code.Writing secure code and dependencies