
MZ
Maike Z\"ufle, Peter Pol\'ak, Sefik Emre Eskimez, Jan Niehues, Peter Bell, Ond\v{r}ej Klejch
· 1 min read
ResearcharXiv cs.CL
Controlling Backchannels in Streamable Full-duplex Models
arXiv:2609.29418v1 Announce Type: new
Abstract: Backchannels, brief acknowledgements like "uh-huh" produced while the other party may still be talking, are central to natural conversation, but full-duplex spoken dialogue models rarely model them explicitly. We introduce a lightweight backchannel head that predicts, from a full-duplex model's own hidden states, when a backchannel should begin. Once this probability crosses a tunable threshold, a backchannel is force-decoded. Attached to both a 7B (PersonaPlex) and a 1B (F-Actor) model, it generalizes across scale. Probing confirms the hidden states anticipate real human timing, and generation evaluation shows more frequent, better-timed backchannels. Human raters judge the resulting backchannels on par with real ones.
Original source
This story was published by arXiv cs.CL and written by Maike Z\"ufle, Peter Pol\'ak, Sefik Emre Eskimez, Jan Niehues, Peter Bell, Ond\v{r}ej Klejch. SyncAI.news shows a preview; the complete article is on the publisher's site.
Read the full story on arxiv.org


