
JH
Jiakai Huang, Zhongbo Wu, Siyu Xu, Zheng Zhang, Zihan Wang, Shan You, Chang Xu, Tao Huang
· 1 min read
ResearcharXiv cs.AI
Foresight Without Seeing: Latent Futures for World Action Models
arXiv:2608.11605v2 Announce Type: replace
Abstract: World Action Models (WAMs) connect visual prediction with robot control, but supplying predictive context often requires expensive future-video generation. Direct policies avoid this cost but lack an explicit interface for accessing future-indexed predictive information. We introduce ForeWAM, a World Action Model that separates forecasting from rendering to expose and shape latent predictive context for efficient control. Its core mechanism, Future-KV, performs a single Video DiT prefill over the current visual latent and noise-initialized future slots, then reuses the resulting key-value states throughout action denoising. To make this context relevant to control, we introduce dynamics registers supervised by latent actions from a frozen teacher during training, encouraging representations of interaction-induced transitions. This reusable context supports a lightweight, single-layer action decoder. We evaluate ForeWAM on LIBERO, LIBERO-Plus, RoboCasa, and real-world manipulation tasks. Without additional policy-level embodied pretraining, ForeWAM improves RoboCasa success by 9.7 percentage points over Fast-WAM at the same budget of 50 demonstrations per task, reaching 59.2%. With a single-layer decoder, it achieves 77.6% success on LIBERO-Plus and reduces policy-query latency to 88.7 ms on an NVIDIA A800, delivering a 6.27-fold speedup over Fast-WAM. These results show that latent predictive computation provides useful foresight for robust, efficient control without explicit future-video generation.
Original source
This story was published by arXiv cs.AI and written by Jiakai Huang, Zhongbo Wu, Siyu Xu, Zheng Zhang, Zihan Wang, Shan You, Chang Xu, Tao Huang. SyncAI.news shows a preview; the complete article is on the publisher's site.
Read the full story on arxiv.org


