SyncAI.news, a Varaisys broadcasting
Phoneme-guided TTS augmentation for ASR: A unified pipeline and multilingual evaluation
ZW

Zhen Wang, TianRui Wu, RongQi Han, Hao Wu, Wei Liang, Wei Xu

· 1 min read

ResearcharXiv cs.CL

Phoneme-guided TTS augmentation for ASR: A unified pipeline and multilingual evaluation

arXiv:2608.26697v3 Announce Type: replace Abstract: Synthetic speech can provide additional supervision for automatic speech recognition (ASR), but constructing useful synthetic training data requires choosing both what to synthesize and how to synthesize it. We present a phoneme-guided text-to-speech (TTS) augmentation pipeline for ASR that connects multilingual speech generation with candidate-text selection and reference-speech quality control. Within this pipeline, we propose phoneme-frequency-guided selection (PFGS), which uses phoneme frequencies from real ASR training transcripts to prioritize candidate texts containing common phonetic content. Experiments with separate monolingual ASR systems cover four languages and 13 test sets. With random text selection, the pipeline improves recognition on 11 test sets at one or more synthesis ratios. PFGS further outperforms random selection on nine test sets, with relative word error rate (WER) reductions of up to 19.3%. An ablation with fixed target texts and synthesis counts further shows the benefit of reference-speech filtering. These results support using real-data phoneme statistics to guide the construction of effective synthetic supervision for ASR.

Original source

This story was published by arXiv cs.CL and written by Zhen Wang, TianRui Wu, RongQi Han, Hao Wu, Wei Liang, Wei Xu. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on arxiv.org

Similar News