retrieval-summarization-task-pulls-definitions-surface-term-match-general
mechanismsingle paperpending review

When retrieval for a summarization task pulls definitions by surface term match from a general encyclopedia, ambiguous terms fetch the wrong sense and the model incorporates that unsupported background into its output, lowering claim-level factuality below the no-retrieval baseline even when the prompt tells it to ignore unrelated retrieved content.

Ingested from a paper but not yet reviewed by a human. It is deliberately inert: it does not move any technique’s standing, does not count toward the backtest, and is excluded anywhere a claim would carry weight. Read the source before relying on it.

Capability: Stating false facts confidently

Observed on

Radiology-report lay summarization with a small general model (Qwen3.5-0.8B) and BioBART-v2-large, retrieval via Wikipedia API first-sentence definitions for terms extracted from t.

Sources

Status: pending-reviewLast checked: 2026-09-09Evidence activity: not checked yet
Contest this claim

Disagreeing is the most useful thing you can do here. Both sides of every contested claim in this catalog were assembled by the same person, which is its weakest point.

Related claims

Notes

Drafted from the paper by a model and filed unreviewed. Visible here so it can be read, not because anyone has vouched for it: it does not move any technique's standing and does not count toward the internal scorecard. Drafted confidence: medium. Falsifier as drafted: Same pipeline with the same surface-form retrieval showing claim-level factuality (FENICE/SummaC) at or above the no-retrieval baseline, or showing that the factuality drop persists when the retrieved definitions are all sense-correct. Drafted stance toward retrieval-moves-the-trust-boundary-rather-than-removing-it: supports -- Retrieved definitions were treated as reliable context and injected unsupported content, so grounding shifted the failure into the retrieval channel rather than removing it.