
YG
Yiwen Guan, Jacob Whitehill
· 1 min read
ResearcharXiv cs.CL
Asymmetric Classifier-Free Guidance for Target-Speaker ASR
arXiv:2609.30476v1 Announce Type: cross
Abstract: Target-speaker automatic speech recognition (TS-ASR) must identify and transcribe a desired speaker under varying overlap and noise conditions. These changes alter the acoustic evidence for the target speaker in the speech mixture, motivating inference-time calibration of speaker conditioning. We introduce asymmetric classifier-free guidance (CFG) for TS-ASR using Whisper: the speaker-conditioned branch predicts the target transcript, while the speaker-unconditioned branch predicts serialized multi-speaker transcripts. CFG adjusts the contribution of speaker conditioning during decoding through a single guidance scale. We select a global guidance scale on target-domain development data and train a lightweight encoder-based predictor to adjust it for each utterance, keeping the recognition model fixed. Under domain shifts, our full system achieves relative word error rate (WER) reductions of up to 21.8% over the condition-only baseline, and 5.6% over standard conditional decoding of the same CFG-trained model. Oracle analysis shows that substantially larger WER reductions are possible through utterance-level scale selection and identifies how beneficial adjustments vary with domain shifts.
Original source
This story was published by arXiv cs.CL and written by Yiwen Guan, Jacob Whitehill. SyncAI.news shows a preview; the complete article is on the publisher's site.
Read the full story on arxiv.org


