training · synthetic-desycophancy-training

Fine-tune on opinion-irrelevant examples

Fine-tune on synthetic prompts where a stated user opinion must not change the answer.

Generate prompts that pair a factual question with a user's stated view, where the correct answer is independent of that view, and fine-tune the model to answer the same way regardless. A small amount of such data reduces sycophancy without hurting other capabilities.

Requires: Requires fine-tuning access.

Addresses: Telling the user what they want to hear

Contexts: Chat assistant

Does it work?

argued, not measured1 supporting · 0 contesting source · last moved 2026-09-04

Efficacy claims — what this technique actually moves, under which conditions, and whether that has been contested.

Code

No repository linked yet. Contribute one.

Sources