Preprint
Jul 2026
SAM3D-Guided Object-Centric Representation Alignment for Vision-Language-Action Models
An object-centric 3D representation alignment framework built upon $\pi_0$, using SAM3D as a frozen 3D teacher to provide target-object 3D priors during training, which enables the policy to internalize target-object 3D information while preserving the original RGB-language-to-action inference pipeline without requiring depth, point clouds, masks, SAM3D, or additional 3D modules at test time.
Zongbo Liu, Shan Jie, Xiaoquan Sun et al.
· 0 citations