training-small-instruction-tuned-models-reinforcement-learning-plus-factual-supervision
mechanismsingle paperpending review

When training small instruction-tuned models with reinforcement learning plus factual supervision, routing each atomic fact's verification score only to the tokens that produced it, and down-weighting verifier judgements that do not change when their key evidence is removed, improves factuality benchmark scores over trajectory-level or reasoning- step-level factual rewards.

Ingested from a paper but not yet reviewed by a human. It is deliberately inert: it does not move any technique’s standing, does not count toward the backtest, and is excluded anywhere a claim would carry weight. Read the source before relying on it.

Capability: Checking claims against evidence

Observed on

3B-parameter instruct models (Qwen2.5-3B-Instruct, Llama-3.2-3B-Instruct), GRPO-style RL on multi-hop QA with Wikipedia evidence snippets, GPT-4o for fact extraction and NLI verifi.

Sources

Status: pending-reviewLast checked: 2026-09-07Evidence activity: not checked yet
Contest this claim

Disagreeing is the most useful thing you can do here. Both sides of every contested claim in this catalog were assembled by the same person, which is its weakest point.

Related claims

Notes

Ingested unreviewed on 2026-09-07 and deliberately inert until a human endorses it: it does not move a technique standing, does not count toward the backtest, and is excluded anywhere a claim would carry weight. Drafted confidence: low. Falsifier as drafted: An ablation or replication where token-level routing and reliability weighting give no gain over step-level factual rewards, or where coarse aggregation matches it once tuned equally. Proposed technique, not yet catalogued: fact-aligned reliability-weighted token credit.