arxiv-2609-00067 · paperDo Multimodal LLMs See Before They Read? Diagnosing Contextual Sycophancy
Created: 2026-08-30 · Ingested: 2026-09-09
https://arxiv.org/abs/2609.00067(opens in a new tab)External text can override conflicting image evidence in multimodal large language models, a failure we call multimodal contextual sycophancy. We introduce a 998-case diagnostic that independently varies visual evidence, commonsense priors, and external text, and probe when this failure arises by moving the information boundary around a context-blind visual witness. On abnormal images paired with Gemini-generated false text, GPT-5.1 scores 7.9% under joint conditioning, 49.7% when the context-blind witness report is scored directly, 63.7% under a matched two-call witness-arbiter pipeline that
In brief
When a caption contradicts an image, the order in which a multimodal model reads matters more than how firmly you tell it to trust its eyes. Withholding the text until after the model has written down what it sees recovers most of the lost accuracy, but only for some models.
The diagnostic has 998 cases: 499 counter-intuitive WHOOPS! images where the visible scene conflicts with commonsense, and 499 normal ImageNet controls, each paired with a single generated sentence that is true, false, or irrelevant. 6 models were tested (GPT-5.1, Gemini 2.5 Pro, Qwen3-VL Instruct and Thinking, Claude Sonnet 4.5, Kimi-K2.5) across 7 inference conditions, scored by a GPT-4o-mini judge.
On abnormal false-text cases GPT-5.1 scores 7.9% under joint conditioning despite visual-priority wording, 49.7% for the context-blind witness alone, 63.7% when the witness sees the text, and 84.2% under System-2 Visual Arbitration. It adopts the false claim in 92.0% of cases jointly, 8.4% under S2VA. Arbitration beats witness-only by 19.7–44.1 points on all 6 models, every paired 95% CI excluding zero.
The design isolates causes reasonably: Leaky Witness is prompt-matched to S2VA, so staged prompting and context withholding are separated. Judge validated against 200 human labels (κ = 0.960).
Boundaries: text helps rather than hurts Claude Sonnet 4.5, Qwen3-Instruct and Kimi-K2.5; true text degrades 3 models by 8.4–13.6 pp. Regenerating the text with GPT-4o lifts GPT-5.1 joint accuracy from 7.9% to 68.0% and makes joint best for 3 of 4 models. 6 models, AI-generated images, 1 sentence of context, no retrieval.
If you deploy retrieval or captions alongside images, calibrate the information boundary per model and per context source rather than assuming isolation helps.
Written from the abstract by claude-opus-5 on 2026-09-09, with every figure checked against it. Not a substitute for the paper.
Referenced by
Claims in this catalog that draw on this source, and whether as support or counterpoint.