training-simple-gameable-environments-generalizes-rarely-tampering-modelsmechanismsingle paper
Training on simple gameable environments generalizes, rarely, to tampering with the model's own reward mechanism.
Capability: Prioritizing safety under conflicting goals
Sources
- Training on simple gameable environments generalizes, rarely, to tampering with the model's own reward mechanism.
Disagreeing is the most useful thing you can do here. Both sides of every contested claim in this catalog were assembled by the same person, which is its weakest point.
Related claims
- Training against an explicit set of principles produces models that are more harmless without becoming evasive.Prioritizing safety under conflicting goals