
CT
Chihiro Taguchi, Yotaro Kubo, Rujikorn Charakorn
· 1 min read
ResearcharXiv cs.AI
BaLEEN: Biasing with Latent Encoded Entities for Context-Aware ASR
arXiv:2609.36913v1 Announce Type: cross
Abstract: Transcribing domain-specific entities and rare proper nouns remains a major challenge in automatic speech recognition (ASR). In this paper, we propose BaLEEN (Biasing with Latent Encoded Entities), a lightweight, hypernetwork-based framework for dynamic contextual adaptation without fine-tuning the underlying ASR model. BaLEEN encodes variable-length contextual keywords using a pretrained language model, compresses them into a fixed sequence of latent vectors via a Perceiver bottleneck, and injects context-dependent bias vectors directly into the intermediate encoder representations of the ASR model. Because both the language model and the backbone ASR model remain entirely frozen during training, BaLEEN operates as a plug-and-play adapter that incurs zero computational overhead at inference time when context biases are precomputed. We evaluate our method on a CTC-based ASR model using a Wikipedia-derived corpus with annotated named entities and synthetic speech. Experimental results demonstrate that BaLEEN reduces keyword miss rate by 8.7% on the test set relative to the unbiased baseline while simultaneously improving overall word error rate by 21% and character error rate by 28%.
Original source
This story was published by arXiv cs.AI and written by Chihiro Taguchi, Yotaro Kubo, Rujikorn Charakorn. SyncAI.news shows a preview; the complete article is on the publisher's site.
Read the full story on arxiv.org


