withholding-the-users-opinion-is-the-no-training-alternativemechanismmechanism reasoning
Sycophancy is triggered by the user's stated view being in context, so the cheapest control is not putting it there — ask for the answer before the opinion, or withhold the opinion entirely. Fine-tuning is the answer for the cases where the opinion has to be in context and the answer still must not move.
Capability: Telling the user what they want to hear · Behavior, Chat assistant
Sources
- Establishes the trigger by construction: models agree with objectively incorrect addition statements when the user endorses them, despite answering correctly without the endorsement. The no-opinion condition is the paper's own baseline.
Disagreeing is the most useful thing you can do here. Both sides of every contested claim in this catalog were assembled by the same person, which is its weakest point.
Related claims
- Fine-tuning on simple synthetic examples where the user's opinion is irrelevant to the answer reduces sycophancy substantially.Telling the user what they want to hear
- When a user profile of stated attributes and preferences is placed in context, models agree with the user far more than with no profile — across 13 open and closed models the profile, not retrieved conversation memory, is the dominant driver, and inverting the stated preference flips almost all responses to the opposite side.Telling the user what they want to hear · unreviewed
- In open-weight instruction-tuned models, interventions that suppress caving to user pushback (DPO, SFT on chosen responses, or activation steering) also tend to reduce the model's rate of correcting a wrong answer when genuine supporting evidence is supplied, because the two answer-flip behaviors run on overlapping MLP neurons and attention heads with positively aligned steering directions.Telling the user what they want to hear · unreviewed
Notes
Backing is mechanism-reasoning, not single-paper, and the gap is specific: the paper's design shows the opinion is what moves the answer, but nobody has tested answer-before-opinion ordering as a deliberate intervention, and I would expect it to leak in multi-turn settings where the view was stated earlier and is still in context. Worth an own-observation claim if I ever run it.