five-assistants-trained-human-feedback-consistently-show-sycophancymechanismsingle paper
Five assistants trained with human feedback consistently show sycophancy across tasks, and human preference data itself rewards it.
Evidence for: Goodhart's law (holds)
Capability: Telling the user what they want to hear
Sources
- Five assistants trained with human feedback consistently show sycophancy across tasks, and human preference data itself rewards it.
Disagreeing is the most useful thing you can do here. Both sides of every contested claim in this catalog were assembled by the same person, which is its weakest point.
Related claims
- Rewarding sycophantic behavior generalizes to more serious specification gaming in a small fraction of cases.Telling the user what they want to hear
- Commercial assistants show a large accuracy drop when the relevant information sits in a long interaction history, especially for updates and multi-session reasoning.Remembering across sessions
- In open-weight instruction-tuned models, interventions that suppress caving to user pushback (DPO, SFT on chosen responses, or activation steering) also tend to reduce the model's rate of correcting a wrong answer when genuine supporting evidence is supplied, because the two answer-flip behaviors run on overlapping MLP neurons and attention heads with positively aligned steering directions.Telling the user what they want to hear · unreviewed
- When preference pairs are built by having one strong model write all chosen responses and one weak model all rejected responses (delta learning), contrastive objectives such as DPO transfer the chosen model's sycophantic-agreement rate to the student even though no individual example contains sycophancy — the student's rate tracks the log-ratio of the two teachers' sycophancy rates, so reversing which teacher is 'chosen' drives sycophancy far below the SFT starting point.Telling the user what they want to hear · unreviewed