grounding-generation-retrieved-documents-improves-factual-accuracy-knowledge-intensivemechanismsingle paper
Grounding generation in retrieved documents improves factual accuracy on knowledge-intensive tasks.
Capability: Stating false facts confidently
Sources
- Grounding generation in retrieved documents improves factual accuracy on knowledge-intensive tasks.
Disagreeing is the most useful thing you can do here. Both sides of every contested claim in this catalog were assembled by the same person, which is its weakest point.
Related claims
- Grounding answers in retrieved documents reduces unsupported generation but transfers the trust problem to the retrieval channel: whatever lands in context is treated as reliable, which shows up both as following instructions planted in retrieved content and as deferring to whichever account is better represented when sources disagree.Stating false facts confidently
- When retrieval for a summarization task pulls definitions by surface term match from a general encyclopedia, ambiguous terms fetch the wrong sense and the model incorporates that unsupported background into its output, lowering claim-level factuality below the no-retrieval baseline even when the prompt tells it to ignore unrelated retrieved content.Stating false facts confidently · unreviewed
- Training the model to decide when to retrieve and to critique whether passages support its output improves factuality and citation accuracy.Checking claims against evidence
- On multi-document question answering with GPT-3.5 Turbo's 16k-context variant, accuracy dropped by more than twenty points when the document containing the answer was moved from the start or end of the context to the middle, with nothing else about the task changed.Losing information in long inputs
- When training small instruction-tuned models with reinforcement learning plus factual supervision, routing each atomic fact's verification score only to the tokens that produced it, and down-weighting verifier judgements that do not change when their key evidence is removed, improves factuality benchmark scores over trajectory-level or reasoning- step-level factual rewards.Checking claims against evidence · unreviewed