arxiv-2608-25662 · paperOverview of SHROOM-Visions 2026: A Shared Task on Hallucination Detection in Large Vision-Language Models
Created: 2026-08-26 · Ingested: 2026-09-09
https://arxiv.org/abs/2608.25662(opens in a new tab)In 2026, we held the fourth iteration of the SHROOM Shared Task series: SHROOM-Visions (\textbf{S}hared-task on \textbf{H}allucinations and \textbf{R}elated \textbf{O}bservable \textbf{O}vergeneration \textbf{M}istakes in \textbf{Vision} language model\textbf{s}), which is hosted at the UncertaiNLP Workshop co-located with EMNLP 2026. Following the success of the 2024 and 2025 tasks, this time we aim to tackle hallucinations through a model-agnostic detection task focused on large vision-language models. Building on the recently introduced SHEEP dataset, designed for long-term evaluation acros
In brief
Fine-grained hallucination detection in vision-language models is still far from solved: after a full shared task with heavy fine-tuning and ensembling, no system exceeded 0.6 on any language or metric, and leaderboard ranks are unstable enough that ordering claims mostly do not survive resampling.
SHROOM-Visions 2026 asked systems to mark character spans in image-conditioned responses and assign 1 of 5 hallucination categories (invention, mischaracterization, OCR problem, miscounting, other) across Chinese, English, French and Italian. It uses the SHEEP dataset: 20,000 samples, LVLM outputs from 5 models plus 1,600 human-written items, with human-written examples appearing only in the test set. 27 teams made 623 submissions; approaches ranged from LoRA-tuned Qwen-VL to text-only XLM-RoBERTa taggers and VLM-judge ensembles.
Best systems averaged 0.58 character-level correlation, 0.46 label-conditioned correlation, 0.51 IoU, 30–40 points above baselines. Most means fell below 0.4. English was hardest; Chinese scored highest.
The uncertainty analysis is the stronger contribution: bootstrap resampling (25,000 samples) put the top English system's mean rank at 2.9 with a 95% interval spanning positions 1 to 10, and rank intervals reached 15 positions in English. Rankings across Random, Human and Silver partitions correlated at ρ 0.575–0.973, with random sampling the most volatile; Friedman tests showed strategy-dependent scores for all metrics in EN and ZH.
Systems abstain far more than the gold data warrants: empty annotations cover 15.5–25.5% of test instances but systems predict empty for 30.7–46.0%, and the metrics do not score abstention. Treat single-number leaderboard positions on span-level hallucination benchmarks as noise unless accompanied by rank intervals.
Written from the abstract by claude-opus-5 on 2026-09-09, with every figure checked against it. Not a substitute for the paper.
Referenced by
Claims in this catalog that draw on this source, and whether as support or counterpoint.