
MH
Meibo Hu, Guohao Sun, Annemarie D. Ross, Sheng Li, Zhiqiang Tao
· 1 min read
ResearcharXiv cs.CV
Attention-Steered Vision-Language Models for Sign Language Translation
arXiv:2608.00235v2 Announce Type: replace
Abstract: Vision-language models (VLMs) have emerged as a powerful framework for multimodal video understanding. However, they remain limited in the sign language translation task, where we identify a key failure mode of existing VLMbased translators: poor spatial-temporal visual grounding. In particular, we find that standard next-token cross-entropy does not directly provide signal for where and when the model should attend, causing models to overlook sign-relevant regions and frames. To address this challenge, we propose AttnSign, a VLM-based spatial-temporal attention steering framework for sign language translation. AttnSign first introduces spatial attention supervision for sign-relevant regions, such as face and hands, in each frame; then develops an RL-based motion-cadence steering method that encourages the model to explore and focus on sign-level keyframes. Experimental results on How2Sign and OpenASL benchmarks show that our proposed AttnSign consistently outperforms existing methods.
Original source
This story was published by arXiv cs.CV and written by Meibo Hu, Guohao Sun, Annemarie D. Ross, Sheng Li, Zhiqiang Tao. SyncAI.news shows a preview; the complete article is on the publisher's site.
Read the full story on arxiv.org


