SyncAI.news, a Varaisys broadcasting
Beyond Speech Captions: Speech-Rewarded Style Planning for Conversational Text-to-Speech
SZ

Shiao Zhu, Lianbo Liu, Sizhen Lyu, Yuzhe Wang, Sheng Li, Takahiro Shinozaki

· 1 min read

ResearcharXiv cs.CL

Beyond Speech Captions: Speech-Rewarded Style Planning for Conversational Text-to-Speech

arXiv:2610.11461v1 Announce Type: cross Abstract: Natural-language style descriptions provide an interpretable interface between large language models (LLMs) and controllable text-to-speech (TTS). However, using descriptions as pseudo-labels compresses target acoustics into text, and descriptive fidelity need not imply effective control of a particular synthesizer. We empirically show that speech-text alignment only weakly predicts downstream acoustic similarity among candidate instructions for the same utterance. We therefore propose Speech-Rewarded Style Planning (SRSP), which trains a text-based style planner through a frozen downstream TTS model. Given dialogue history and response text, the planner generates candidate instructions and is optimized with group-relative policy optimization (GRPO), using the teacher-forced likelihood of target speech tokens as the reward. On an English subset of the ISCSLP 2026 CoT-TTS corpus, SRSP achieves higher speech-style and emotion similarity to target speech and lower mel-cepstral distortion than the Base LLM and target-audio-informed captioning baselines. LLM-based expressive speech evaluation further shows gains over all baselines in contextual appropriateness and reference consistency.

Original source

This story was published by arXiv cs.CL and written by Shiao Zhu, Lianbo Liu, Sizhen Lyu, Yuzhe Wang, Sheng Li, Takahiro Shinozaki. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on arxiv.org

Similar News