general-purpose-factuality-verifiers-entailment-consistency-rag
observationsingle paperpending review

General-purpose factuality verifiers (entailment-, consistency-, and RAG-trained) and zero-shot frontier LLM judges perform near chance at spotting unsupported claims in scientific peer reviews that must be checked against the full submitted paper, because they cannot separate paper-unsupported assertions from legitimate evaluative critique; small models fine-tuned on in-domain examples beat them by a wide margin.

Ingested from a paper but not yet reviewed by a human. It is deliberately inert: it does not move any technique’s standing, does not count toward the backtest, and is excluded anywhere a claim would carry weight. Read the source before relying on it.

Capability: Stating false facts confidently

Observed on

Peer reviews of ML conference papers (ICLR/NeurIPS on OpenReview) with synthetically injected hallucinations; detection at review and sentence level with BM25 or full-paper evidenc.

Sources

  • Measured on 12K papers and 38K reviews; four specialized verifiers gave review-level MCC of at most 0.03, prompted LLMs including GPT-5.2 stayed at 0.15-0.25 review-level MCC, while a fine-tuned Qwen2.5-3B reached 0.69 and Qwen3-32B 0.81. Hallucinations are LLM-injected rather than naturally occurring, so difficulty may reflect the injection distribution; a small authentic-review check (20 annotated cases) showed the fine-tuned detector recovered all of them at 22% false positive rate.
Status: pending-reviewLast checked: 2026-09-09Evidence activity: not checked yet
Contest this claim

Disagreeing is the most useful thing you can do here. Both sides of every contested claim in this catalog were assembled by the same person, which is its weakest point.

Related claims

Notes

Drafted from the paper by a model and filed unreviewed. Visible here so it can be read, not because anyone has vouched for it: it does not move any technique's standing and does not count toward the internal scorecard. Drafted confidence: medium. Falsifier as drafted: An off-the-shelf verifier or zero-shot frontier LLM reaching review-level MCC comparable to the in-domain fine-tuned models on this or a similar peer-review grounding benchmark. Proposed technique, not catalogued: Taxonomy-driven hallucination injection for review verification benchmarks.