
RR
Rayhan Rashed, Senja Filipi, Ross Cutler
· 1 min read
ResearcharXiv cs.CV
Adapting Personalized Speech Enhancement for Low-Latency Audio-Visual Target-Speaker Extraction
arXiv:2609.30631v1 Announce Type: cross
Abstract: Online audio-visual target-speaker extraction aims to remove competing voices while preserving speech quality and bounding lookahead. Existing extractors are built and evaluated for separation on synthetic mixtures, leaving listening quality and meeting behavior largely untested. We introduce Audio-Visual Personalized Voice Quality Enhancement (AV-PVQE), which approaches these requirements from the other direction. We start from a personalized speech enhancement model that reconstructs a requested voice at high quality but confuses the target in 46% of two-speaker mixtures despite clean enrollment. Adding mouth features at its speaker-conditioning input and jointly fine-tuning the visual and reconstruction networks reduces this rate to 1.6%, with no future frames and 20 ms of algorithmic delay. Compared with an online autoregressive audio-visual extractor, AV-PVQE yields separation gains on two synthetic benchmarks and larger gains on recorded meetings, and keeps its advantage on excerpts with more speakers than the fine-tuning mixtures. In personalized P.835 listening tests on two meeting corpora, it improves overall quality over this extractor by 0.57 and 0.63 MOS, with similar mean rating relative to the starting model. Preservation and rejection tests show that it keeps the target intact when no competing voice is present and suppresses competing speech when the target is absent.
Original source
This story was published by arXiv cs.CV and written by Rayhan Rashed, Senja Filipi, Ross Cutler. SyncAI.news shows a preview; the complete article is on the publisher's site.
Read the full story on arxiv.org


