questions-whose-premises-false-or-internally-contradictory-prompting-mid-size
mechanismsingle paperpending review

On questions whose premises are false or internally contradictory, prompting mid-size instruct models (Qwen2.5-7B, LLaMA-3.1-8B, Gemma3-12B) to reason step by step can lower accuracy below plain answering, because the reasoning chain elaborates from the flawed premise instead of challenging it.

Ingested from a paper but not yet reviewed by a human. It is deliberately inert: it does not move any technique’s standing, does not count toward the backtest, and is excluded anywhere a claim would carry weight. Read the source before relying on it.

Capability: Stating false facts confidently

Observed on

Mid-size open instruct models, short-form QA with planted false premises or contradictions (TruthfulQA multiple-choice, FalseQA, the authors' MisFactQA); not tested on premise-clea.

Sources

  • Measured on three open instruct models across three datasets against a default-prompt baseline; CoT fell below Original on TruthfulQA for all three models and on FalseQA for LLaMA and Gemma, and also below Original for GPT-4o-mini and DeepSeek-V3. Judged by an o3-mini automated judge with reported human agreement. Single paper; the causal story (reasoning amplifies the accepted premise) is argued from case studies rather than isolated experimentally.
Status: pending-reviewLast checked: 2026-09-07Evidence activity: not checked yet
Contest this claim

Disagreeing is the most useful thing you can do here. Both sides of every contested claim in this catalog were assembled by the same person, which is its weakest point.

Related claims

Notes

Ingested unreviewed on 2026-09-07 and deliberately inert until a human endorses it: it does not move a technique standing, does not count toward the backtest, and is excluded anywhere a claim would carry weight. Drafted confidence: medium. Falsifier as drafted: On false- premise question sets, CoT prompting matching or beating direct answering across model families, with the drop not reproducing. Proposed technique, not yet catalogued: Verify input premises before answering.