7-8b-vision-language-models-fine-tuning-visual-text-embeddings-toward-perturbation-averaged
observationsingle paperpending review

In 7-8B vision-language models, fine-tuning visual and text embeddings toward perturbation-averaged and ground-truth anchors before cross-modal contrastive alignment lowers object and attribute hallucination rates on captioning and existence benchmarks without reducing object coverage or general VQA accuracy.

Ingested from a paper but not yet reviewed by a human. It is deliberately inert: it does not move any technique’s standing, does not count toward the backtest, and is excluded anywhere a claim would carry weight. Read the source before relying on it.

Capability: Grounding answers in the image

Observed on

Single-image, single-turn settings at 7-8B scale (LLaVA-1.5, LLaVA-1.6, Qwen3-VL-8B); 2K training pairs drawn from RLHF-V chosen/rejected responses; evaluated on AMBER, ObjHal, MMH.

Sources

  • Measured on three backbones against re-run base models and re-run DPO and CHiP under matched data; AMBER CHAIR falls 7.8->4.2, 8.3->3.1, 5.9->2.9, with coverage rising and VQA-v2/TextVQA matched or improved. Ablations separate the intra-modal stage from the alignment stage. Not isolated: no ablation against a data-matched plain SFT baseline, and some comparison rows are literature-reported rather than controlled.
Status: pending-reviewLast checked: 2026-09-09Evidence activity: not checked yet
Contest this claim

Disagreeing is the most useful thing you can do here. Both sides of every contested claim in this catalog were assembled by the same person, which is its weakest point.

Related claims

Notes

Drafted from the paper by a model and filed unreviewed. Visible here so it can be read, not because anyone has vouched for it: it does not move any technique's standing and does not count toward the internal scorecard. Drafted confidence: low. Falsifier as drafted: Replication on the same backbones showing that hallucination reductions are matched by drops in object coverage or in VQA-v2/TextVQA accuracy, or that a plain generation-loss fine-tune on the same 2K data gives the same hallucination reduction without the stabilization objectives. Proposed technique, not catalogued: Perturbation-invariance fine-tuning for vision-language embeddings.