SyncAI.news, a Varaisys broadcasting
Supervise What Decides Success: Criterion-Aligned Auxiliary Losses for Latent World-Model Planning
TH

Takumi Hara, Kanata Suzuki

· 1 min read

ResearcharXiv cs.LG

Supervise What Decides Success: Criterion-Aligned Auxiliary Losses for Latent World-Model Planning

arXiv:2610.01224v1 Announce Type: new Abstract: Latent world models plan by scoring candidate action sequences with distances in latent space. However, task success is judged by physical quantities, which we call the success-criterion quantities. In all four latent world models we examine, the end-effector position is encoded in the latent state with an error larger than the success criterion allows. Such a latent state cannot separate successful candidates from failing ones. We propose an auxiliary loss that uses success-criterion quantities as training targets, whereas existing latent world models take them only as inputs. During training, a linear head on the encoder and predictor outputs regresses the success-criterion quantities, and the regression error is added to the training loss. The head is discarded after training, so the model, its cost, and its inputs at test time are unchanged. This loss alone improves the success rate on PushT and cube by 3.5% and 3.4% (absolute), respectively, and both improvements are statistically significant. A success criterion thus specifies what a world model must retain in its latent state, and we show that it can serve directly as a training target.

Original source

This story was published by arXiv cs.LG and written by Takumi Hara, Kanata Suzuki. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on arxiv.org

Similar News