SyncAI.news, a Varaisys broadcasting
DriftTTS: Few-Step Text-to-Speech Without Distillation via Distribution-Matching Drift
MN

Mohammad Nur Hossain Khan, Subrata Biswas, Bashima Islam

· 1 min read

ResearcharXiv cs.AI

DriftTTS: Few-Step Text-to-Speech Without Distillation via Distribution-Matching Drift

arXiv:2610.03390v1 Announce Type: cross Abstract: Few-step neural text-to-speech models often rely on short- ened diffusion or flow-matching schedules, or on distillation from pretrained multi-step teachers. To avoid these depen- dencies, we present DriftTTS, a few-step mel-spectrogram generator trained without a generative teacher, distillation, or adversarial discrimination. DriftTTS uses a distribution- matching drift objective in a mel-domain feature space defined by raw mels and a frozen masked-autoencoder encoder pretrained on the same LJSpeech training split. On-policy rollout trains the decoder on its own interme- diate states and supports inference up to the trained roll- out depth. On LJSpeech, DriftTTS at NFE=4 achieves 3.87 dB MCD and 3.7% WER, compared with 3.85 dB and 3.4% for Matcha-TTS. In a fully paired blind listen- ing test, DriftTTS obtains 4.18 MOS, compared with 3.96 for Matcha-TTS and 4.22 for ground truth. These results demonstrate competitive few-step synthesis without a pre- trained generative teacher. Code can be found at https: //github.com/BASHLab/driftTTS.git

Original source

This story was published by arXiv cs.AI and written by Mohammad Nur Hossain Khan, Subrata Biswas, Bashima Islam. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on arxiv.org

Similar News