SyncAI.news, a Varaisys broadcasting
Repetition, Not Length: Isolating the Counting Failure in Neural Text-to-Speech
KB

Kirill Borodin, Vasilii Kudryavtsev, Maxim Maslov, Grach Mkrtchian

· 1 min read

ResearcharXiv cs.CL

Repetition, Not Length: Isolating the Counting Failure in Neural Text-to-Speech

arXiv:2609.36974v1 Announce Type: new Abstract: Text-to-speech models loop, truncate and lose count on text that repeats a phrase many times. We show that repetition itself is what breaks them, not the length that comes with it. Every repeated sentence in our test set is paired with a control of matched sentence and word count in which no word ever repeats back-to-back. Six models from three architectures render the controls almost perfectly and fail the repeated twins: 94.3% against 18.2% exactly right at k >= 6. The gap survives greedy decoding, repetition-penalty sweeps, four independent speech recognisers and 420 analysis specifications without once reversing sign; a held-out fourth architecture lands within a point of its predicted gap, and one of two non-autoregressive baselines shows the same failure. Varying the period of the text shows the failure grows smoothly with periodicity, half of it surviving when no word is adjacent to itself.

Original source

This story was published by arXiv cs.CL and written by Kirill Borodin, Vasilii Kudryavtsev, Maxim Maslov, Grach Mkrtchian. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on arxiv.org

Similar News