
PW
Peidong Wang, Jian Xue, Jinyu Li
· 1 min read
ResearcharXiv cs.AI
LOGIC: Efficient and Robust Contextual Biasing for Speech LLMs via Logit-Space Integration
arXiv:2601.15397v4 Announce Type: replace
Abstract: Recognizing entity phrases remains a critical challenge for speech large language models. Existing prompting methods lack an explicit decoding-time biasing weight, limiting their controllability. Generative error correction methods can introduce hallucinated over-corrections. To address these limitations, we propose LOGIC (logit-space integration for contextual biasing), a robust framework operating directly in the logit space. By decoupling context injection from input processing, LOGIC enables explicit control over the biasing strength. Extensive experiments with an open-source speech large language model across 11 locales demonstrate that LOGIC achieves an average 9% relative reduction in entity word error rate, with an average false alarm rate increase of 0.3% and a 2.8% relative runtime overhead. When combined with prompting, LOGIC can reduce entity word error rate by 5% relative to the prompt-only method.
Original source
This story was published by arXiv cs.AI and written by Peidong Wang, Jian Xue, Jinyu Li. SyncAI.news shows a preview; the complete article is on the publisher's site.
Read the full story on arxiv.org


