fine-tuning-base-model-instruction-data-removing-training-targets-any
mechanismsingle paperpending review

When fine-tuning a base model on instruction data, removing from the training targets any factual claim the base model cannot consistently recall raises the supported-claim rate of later long-form generations, but the gain comes largely from a more conservative response policy — more refusals and fewer supported claims per answer — not from generating more correct facts at equal coverage.

Ingested from a paper but not yet reviewed by a human. It is deliberately inert: it does not move any technique’s standing, does not count toward the backtest, and is excluded anywhere a claim would carry weight. Read the source before relying on it.

Capability: Stating false facts confidently

Observed on

Shown on Qwen3-4B-Base and OLMo 3 7B fine-tuned on the small English first-turn OASST1 set, evaluated on entity-centric long-form factuality benchmarks (WildHalu, Biography) with a.

Sources

  • Compared four knowledge-alignment methods against standard SFT under one training setup, plus an ablation varying only the share of known claims (100/50/0%) with prompt count and refusal count fixed, which is the part that isolates causality. Factuality metrics count refusals as fully supported, so the headline gain is partly definitional; the authors say so and report the falling supported-claim counts. Single dataset, small SFT corpus, automatic verifiers.
Status: pending-reviewLast checked: 2026-09-09Evidence activity: not checked yet
Contest this claim

Disagreeing is the most useful thing you can do here. Both sides of every contested claim in this catalog were assembled by the same person, which is its weakest point.

Related claims

Notes

Drafted from the paper by a model and filed unreviewed. Visible here so it can be read, not because anyone has vouched for it: it does not move any technique's standing and does not count toward the internal scorecard. Drafted confidence: medium. Falsifier as drafted: A run where recall-filtered targets raise supported-claim percentage while holding refusal count and number of supported claims per response equal to standard SFT would show the gain is not coverage-driven; conversely, showing no factuality gain at all when the share of known claims in targets is varied would falsify the mechanism. Proposed technique, not catalogued: Recall-consistency filtering of fine-tuning targets.