training ·
synthetic-desycophancy-trainingFine-tune on opinion-irrelevant examples
Fine-tune on synthetic prompts where a stated user opinion must not change the answer.
Generate prompts that pair a factual question with a user's stated view, where the correct answer is independent of that view, and fine-tune the model to answer the same way regardless. A small amount of such data reduces sycophancy without hurting other capabilities.
Requires: Requires fine-tuning access.
Contexts: Chat assistant
Does it work?
argued, not measured1 supporting · 0 contesting source · last moved 2026-09-04
Efficacy claims — what this technique actually moves, under which conditions, and whether that has been contested.
Code
No repository linked yet. Contribute one.