open-source-vision-language-models-llava-15-qwen-vl-chat-qwen25-vl-llava-next-training-free-decoding
observationsingle paperpending review

In open-source vision-language models (LLaVA-1.5, Qwen-VL-Chat, Qwen2.5-VL, LLaVA-NeXT), a training-free decoding intervention that detects large layer-to-layer and step-to-step shifts in hidden states and nudges the diverging states back toward their prior value lowers object-hallucination rates, but the lower rate comes with fewer objects mentioned and shorter outputs.

Ingested from a paper but not yet reviewed by a human. It is deliberately inert: it does not move any technique’s standing, does not count toward the backtest, and is excluded anywhere a claim would carry weight. Read the source before relying on it.

Capability: Grounding answers in the image

Observed on

Open-source image-input LVLMs at 3B-7B scale, standard hallucination benchmarks (POPE, CHAIR, HallusionBench); not tested on closed models, video, or multi-image input..

Sources

  • Measured on nine benchmarks across six backbones against vanilla decoding and training-free baselines (VCD, ICD, MemVR, OPERA). CHAIR shows Cs 47.6->28.2 and Ci 13.3->7.3 on LLaVA-1.5 with recall 80.6->72.0 and length 99.7->88.9. Authors report a length-normalized analysis in an appendix not included here, so the informativeness confound is claimed to be addressed but not verifiable from the given text. Single group, fixed thresholds, no seed variance reported.
Status: pending-reviewLast checked: 2026-09-09Evidence activity: not checked yet
Contest this claim

Disagreeing is the most useful thing you can do here. Both sides of every contested claim in this catalog were assembled by the same person, which is its weakest point.

Related claims

Notes

Drafted from the paper by a model and filed unreviewed. Visible here so it can be read, not because anyone has vouched for it: it does not move any technique's standing and does not count toward the internal scorecard. Drafted confidence: medium. Falsifier as drafted: A replication where the same compensation leaves CHAIR recall and response length unchanged while hallucination rates still fall, or where hallucination rates do not fall at all on a held-out backbone. Drafted stance toward training-free-inference-time-hallucination-mitigation-7b-vision-language-models-reported-gains: supports -- On CHAIR the method cuts sentence- and instance-level hallucination rates while recall and average response length both fall, the same coupling of lower hallucination with reduced informativeness. Proposed technique, not catalogued: Hidden-state divergence compensation at decoding time.