llm-based-asr-steering-decoding-toward-acoustic-evidence-only-helps
mechanismsingle paperpending review

In LLM-based ASR, steering decoding toward acoustic evidence only helps if candidates are first restricted to tokens within a small likelihood margin of the greedy token: unconstrained selection by an acoustic compatibility score removes hallucinated outputs but destroys transcription accuracy through unbounded insertions, while the constrained version removes a large share of hallucinations at near-zero WER/CER cost.

Ingested from a paper but not yet reviewed by a human. It is deliberately inert: it does not move any technique’s standing, does not count toward the backtest, and is excluded anywhere a claim would carry weight. Read the source before relying on it.

Capability: Stating false facts confidently

Observed on

Four open LLM-based ASR systems (Qwen3-ASR-1.7B, Qwen2-Audio-7B Base, Kimi-Audio-7B-Instruct, GLM-ASR-Nano) on two 500-utterance stress suites built to elicit code-switching and in.

Sources

  • Paired greedy-vs-LCAR decoding on 8800 frozen detector-positive records plus negative controls, with an LLM detector whose precision was human-audited at 93.3% on a 400-case stratified sample. Ablation on 800 Qwen3 positives isolates the likelihood constraint: acoustic- only selection gives comparable overall hallucination fixing but >600/1000 point WER/CER increases. Normal-speech effects measured on LibriSpeech test-clean and AISHELL2. Gains vary by model and suite; [truncated]
Status: pending-reviewLast checked: 2026-09-09Evidence activity: not checked yet
Contest this claim

Disagreeing is the most useful thing you can do here. Both sides of every contested claim in this catalog were assembled by the same person, which is its weakest point.

Related claims

Notes

Drafted from the paper by a model and filed unreviewed. Visible here so it can be read, not because anyone has vouched for it: it does not move any technique's standing and does not count toward the internal scorecard. Drafted confidence: medium. Falsifier as drafted: An unconstrained acoustic-reranking decoder that matches or beats the likelihood-constrained version on both hallucination removal and standard-test WER/CER, or a constrained variant whose hallucination reduction comes with large WER regressions on normal speech. Proposed technique, not catalogued: likelihood-constrained acoustic reranking.