Skip to content

Improving Cross-embodiment Transfer in Latent Action Models with Action-Similarity Supervision

Sep 2026 · 0 citations · 33 references
Computer Science

TL;DR

This work evaluates cross-embodiment transfer on RoboTwin 2.0 in a controlled setup, two bimanual robots demonstrate disjoint task sets, a policy is trained on all the demonstrations, and each robot is evaluated closed-loop on the tasks only the other demonstrated.

Abstract

As generalist robot policies gain vision and language from web-scale pretraining, demonstrations remain costly to collect and tied to the robot that recorded them. Latent action models (LAMs) address both by learning latent actions from action-free videos that can be shared across embodiments, however, in practice, LAMs are sensitive to background visual noise, and the same motion from two different robots may be encoded with different latents. One solution to the background visual noise is to add an auxiliary loss predicting the robot action from the latent action, further associating the latent action space to the embodiment specific robot action space. We study a different use of the same labels, through action-similarity supervision. The similarity between any two latent actions is trained to match the similarity of the two ground-truth robot action sequences. The ground-truth actions are never predicted by the LAM, so the latent action does not need to encode embodiment specifics. We evaluate cross-embodiment transfer on RoboTwin 2.0 in a controlled setup, two bimanual robots demonstrate disjoint task sets, a policy is trained on all the demonstrations, and each robot is evaluated closed-loop on the tasks only the other demonstrated. With the policy architecture and its hyperparameters, the dataset, and the evaluation protocol fixed, predicting latent actions instead of ground-truth actions more than doubles cross-embodiment success. Given the same ground-truth actions, similarity supervision transfers better than an auxiliary loss that predicts the ground-truth action during the LAM training. Computing the similarities on end-effector motion rather than joint-space motion, and letting the loss compare latent actions across the two robots, gives the best approach of the study.

View source

Similar papers

Preprint Aug 2026

LD4WAM: Learning Latent Dynamics from Human Videos for World Action Models

LD4WAM is presented, which pairs a Latent Dynamics Model trained with semantic reconstruction and real motion alignment with a World Dynamics Action Model built as a mixture-of-transformers (MoT), which preserves full future-video generation and uses learnable queries to distill these latent dynamics from generated fut...

Zhen Shen, Jia-Qi Liang, Jasper Lu et al. · 1 citation
Preprint Aug 2026

JoyAI-RA 0.5: Scaling Robot Manipulation Learning via Dual Action Alignment

JoyAI-RA 0.5 is proposed, a generalist Vision-Language-World-Action framework that couples physical world-dynamics priors with visual semantics and scales manipulation learning across such data via dual action alignment, suggesting that abundant but weakly labeled human experience can be converted into a transferable t...

JoyAI-RA Team · 3 citations · ⚡1
Preprint Aug 2026

SLIM-0.5B: Learning Action-Grounded Predictive Latents for Robot Manipulation

SLIM (Self-supervised Latent Interaction Model), a compact 0.5B-parameter latent interaction policy, which matches or exceeds representative large-scale VLA and world-action-model baselines with fewer parameters, no additional embodied pretraining, lower inference latency, and substantially lower GPU memory usage.

Jing-Kai Wang, Zihan Tang, Gu Zhang et al. · 0 citations
Preprint Sep 2026

UMI-Bridge: Action-Anchored Latent Alignment across Human and Robot Manipulation Data

UMI-Bridge, which uses UMI as an intermediate domain to align representations according to action equivalence rather than pixel similarity, supports action-anchored latent alignment for data-efficient robot learning and UMI-to-robot task transfer.

Hai-Yi Liu, Jin-Ming Ma, Ke Rui et al. · 0 citations
Preprint Sep 2026

GALA: Geometry-Aware Latent Action Modeling for Vision-Language-Action Model Pretraining across Embodiments

Learning large-scale vision-language-action (VLA) models from multi-embodiment datasets remains challenging due to heterogeneous action spaces across end effectors. Although latent action models (LAMs) can learn embodiment-agnostic action representations from diverse video data, existing image-based LAMs often fail to...

Yi-Chen Liu, Pu-Zhen Yuan, Xiang-Pei Zhu et al. · 0 citations
Preprint Sep 2026

CrossBFM: Distilling a Shared Latent Behavior Space Across Humanoid Embodiments

Behavior Foundation Models (BFMs) give humanoids a promptable policy over a latent behavior space, enabling one single vector to represent a motion to imitate, a pose to reach, or a reward to maximize. Forward-Backward representations successfully produce such spaces, but at the cost of hundreds of GPU-hours for a sing...

Tan-Dzung Do, Tuan Dat Phuong, Nico Bohlinger et al. · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.