
YL
Yuxiang Liu, Lizhi Yang, Fengze Xie, Aaron Ames, Yisong Yue
· 1 min read
ResearcharXiv cs.AI
The Layer Mystery of VLA: An Information-Theoretical Analysis of VLA Latent Interface
arXiv:2609.36118v1 Announce Type: new
Abstract: Vision-language-action (VLA) policies connect a pretrained vision-language backbone to an action head through a latent interface, but which backbone layers this interface should expose remains unclear. We study single-layer selection and multi-layer fusion for frozen backbones across three pretrained models and two manipulation benchmarks, LIBERO and CALVIN, with three policy-training seeds per configuration. Across three fusion mechanisms and three layer-subset strategies, 47 of 54 configurations underperform the best observed single-layer policy. Our stastical analysis further confirms that fusion's advantage is very limited. However, the best layer varies substantially across backbones and benchmarks, making layer selection consequential and exhaustive policy sweeps expensive. We further derive a reweighting equivalence between the proposed information-bottleneck objectives for action-conditioned InfoNCE and action prediction, motivating InfoNCE as a proxy for layer quality. Empirically, InfoNCE provides the most consistent positive association with policy success among four evaluated proxies. Selecting the layer with the highest InfoNCE score requires 9-33 times less GPU compute than exhaustive policy sweeps and reduces mean selection regret from 17.89 percentage points for deepest-layer selection to 3.71 points across six settings. Its mean regret is close to the 3.28-3.50 points achieved by fixed-layer heuristics optimized retrospectively using all six oracle sweeps, without requiring closed-loop evaluations during selection.
Original source
This story was published by arXiv cs.AI and written by Yuxiang Liu, Lizhi Yang, Fengze Xie, Aaron Ames, Yisong Yue. SyncAI.news shows a preview; the complete article is on the publisher's site.
Read the full story on arxiv.org


