training-simple-gameable-environments-generalizes-rarely-tampering-models
mechanismsingle paper

Training on simple gameable environments generalizes, rarely, to tampering with the model's own reward mechanism.

Capability: Prioritizing safety under conflicting goals

Sources

Status: activeLast checked: 2026-09-03Evidence activityHow much the field cites the sources under this claimheavily cited in the last 12 months89 in 12mo · 158 total — Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models
Contest this claim

Disagreeing is the most useful thing you can do here. Both sides of every contested claim in this catalog were assembled by the same person, which is its weakest point.

Related claims