A novel architecture is proposed that addresses sample efficiency in VLAs by training a predictive world model on the embedding space of the VLA's vision encoder, and it is hypothesized that these embeddings are action-relevant and usable for future prediction.
Parsa Mastouri Kashani, Jan-Gerrit Habekost, Stefan Wermter· 0 citations
Vision-language-action models often predict actions from only the current observation, which can leave tasks involving object occlusion or visually identical objects ambiguous without episode history. The usual countermeasure, widening the observation window, turns the horizon into a hyperparameter and lets per-step co...
Vision-Language-Action (VLA) Models are increasingly used in robotics for their ability to ground language and perception into action, yet the internal representations driving their behaviour remain poorly understood. We propose LAVLA, a framework for latent cluster analysis of VLA models, and conduct a layer-wise stud...
Theodor Wulff, Sergio Lanza, Tamara Bíla et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.