SyncAI.news, a Varaisys broadcasting
VISTA: Video-Injected Stylized Text-to-Animation
MP

Monseej Purkayastha, Anindita Ghosh, Philipp Slusallek

· 1 min read

ResearcharXiv cs.CV

VISTA: Video-Injected Stylized Text-to-Animation

arXiv:2609.23817v1 Announce Type: new Abstract: We present VISTA, a two-stage framework for generating stylized 3D human motion by fusing structural content from text prompts with expressive style from reference videos, without requiring jointly paired (text, video, stylized motion) triplets. A Dual-channel Autoencoder first maps motion sequences and video clips into a shared latent manifold. A masked autoregressive diffusion backbone then operates within this manifold, injecting video-derived style through a dedicated late-fusion Dual-AdaLN pathway while preserving text-conditioned content structure. A cross-batch unpaired training protocol with latent cycle consistency enables joint learning across separate semantically rich and stylistically diverse datasets. As a proof-of-concept for controllable animation synthesis, we validate VISTA on rendered motion-capture references: it achieves the highest style recognition accuracy among video-conditioned methods while preserving competitive content alignment, and its decomposed 3-way classifier-free guidance provides independent, user-controllable calibration of the content--style balance at inference time.

Original source

This story was published by arXiv cs.CV and written by Monseej Purkayastha, Anindita Ghosh, Philipp Slusallek. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on arxiv.org

Similar News