fine-tuning-simple-synthetic-examples-where-users-opinion-irrelevantmechanismsingle paper
Fine-tuning on simple synthetic examples where the user's opinion is irrelevant to the answer reduces sycophancy substantially.
Capability: Telling the user what they want to hear
Sources
- Fine-tuning on simple synthetic examples where the user's opinion is irrelevant to the answer reduces sycophancy substantially.
Disagreeing is the most useful thing you can do here. Both sides of every contested claim in this catalog were assembled by the same person, which is its weakest point.
Related claims
- Sycophancy is triggered by the user's stated view being in context, so the cheapest control is not putting it there — ask for the answer before the opinion, or withhold the opinion entirely. Fine-tuning is the answer for the cases where the opinion has to be in context and the answer still must not move.Telling the user what they want to hear
- In open-weight instruction-tuned models, interventions that suppress caving to user pushback (DPO, SFT on chosen responses, or activation steering) also tend to reduce the model's rate of correcting a wrong answer when genuine supporting evidence is supplied, because the two answer-flip behaviors run on overlapping MLP neurons and attention heads with positively aligned steering directions.Telling the user what they want to hear · unreviewed
- Rewarding sycophantic behavior generalizes to more serious specification gaming in a small fraction of cases.Telling the user what they want to hear