
PJ
Payam Jome Yazdian, Zoe Stanley, Angelica Lim
· 1 min read
ResearcharXiv cs.CV
SalsaAgent: A multimodal embodied language model for interactive dance generation
arXiv:2605.29219v3 Announce Type: replace
Abstract: Embodied interaction with humanoids depends on bidirectional nonverbal reactivity, coordination, and synchrony to convey cues and move with a partner. For socially interactive embodied agents, reactive motion generation requires expressive full-body motion that remains contextually appropriate while maintaining spatial and temporal synchrony. We present SalsaAgent, a language model that generates expressive, full-body salsa follower motions in reaction to a human leader and music. We formulate partner interaction as nonverbal token passing, extending the vocabulary of a large language model (LLM) to process discrete motion tokens, pairwise relation tokens, and audio tokens. Our method introduces full-body and pairwise-relation tokenizers, aligns language and motion tokens with automatically derived text descriptions of skeleton dynamics, and applies a two-stage token-to-diffusion pipeline. Subjective and objective evaluations show improved motion quality, two-person spatial coordination, and music and partner coordination relative to prior baselines.
Original source
This story was published by arXiv cs.CV and written by Payam Jome Yazdian, Zoe Stanley, Angelica Lim. SyncAI.news shows a preview; the complete article is on the publisher's site.
Read the full story on arxiv.org


