SyncAI.news, a Varaisys broadcasting
HearInContext: A Benchmark for Implicit Context in Speech Recognition
YG

Yifan Gao, Yao Tian, Hongbin Suo

· 1 min read

ResearcharXiv cs.CL

HearInContext: A Benchmark for Implicit Context in Speech Recognition

arXiv:2609.18680v2 Announce Type: replace Abstract: Contextual ASR can benefit from semantic cues or from target words explicitly provided in the context. We introduce HearInContext, a Mandarin-English benchmark that pairs shared synthetic speech with assistant replies supporting different interpretations. The benchmark comprises 3,764 semantic test cases built around homophones. Implicit contexts exclude candidate words; explicit contexts name the target. No-context and unrelated-context controls measure the benefit of relevant history and sensitivity to irrelevant history. Context-capable models benefit from implicit cues but achieve higher target recall with explicit hints. Fine-tuning Qwen3-ASR-1.7B improves implicit-context target recall by 11.4 percentage points in both Mandarin and English, while absolute CER/WER changes on AISHELL-1 and LibriSpeech remain below 0.1 percentage points. Gains extend to explicit conditions excluded from fine-tuning and to Mandarin hotword recognition on real recordings. Code and data are available at https://github.com/OPPO-Mente-Lab/HearInContext

Original source

This story was published by arXiv cs.CL and written by Yifan Gao, Yao Tian, Hongbin Suo. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on arxiv.org

Similar News