span-level-hallucination-detection-leaderboards-where-absolute-scores-low-bootstrap-resampling
observationsingle paperpending review

On span-level hallucination-detection leaderboards where absolute scores are low, bootstrap resampling of the test set moves top systems across wide rank intervals, so point-estimate ordering does not establish that one detector beats another.

Ingested from a paper but not yet reviewed by a human. It is deliberately inert: it does not move any technique’s standing, does not count toward the backtest, and is excluded anywhere a claim would carry weight. Read the source before relying on it.

Capability: Whether the measurement made the finding

Observed on

Shared task on character-level hallucination span detection in vision-language outputs, four languages, 27 teams, 600+ submissions; best average scores around 0.58 correlation and .

Sources

  • Measured: 25,000-sample bootstrap rank distributions per language and metric. Top English system had mean rank 2.9 with a 95% interval spanning ranks 1-10; intervals up to 15 positions in EN, ~5-6 in ZH. Rank correlations across data-construction partitions were also below 1, lowest for randomly sampled examples. Single task, single dataset; the inverse relation between score level and rank stability is observed, not isolated by design.
Status: pending-reviewLast checked: 2026-09-09Evidence activity: not checked yet
Contest this claim

Disagreeing is the most useful thing you can do here. Both sides of every contested claim in this catalog were assembled by the same person, which is its weakest point.

Related claims

Notes

Drafted from the paper by a model and filed unreviewed. Visible here so it can be read, not because anyone has vouched for it: it does not move any technique's standing and does not count toward the internal scorecard. Drafted confidence: medium. Falsifier as drafted: Bootstrap resampling of the same test set yields narrow rank intervals (e.g. top systems confined to one or two positions) despite the low absolute scores. Proposed technique, not catalogued: Bootstrap rank intervals reported alongside leaderboard scores.