Skip to content
Preprint

Proprioception-Anchored Cross-Modal Pretraining for Zero-Shot Sim-to-Real Contact-Rich Assembly

Sep 2026 · 0 citations · 48 references
Computer Science

TL;DR

PACE (Proprioception-Anchored Cross-Modal Encoder) is presented, which supervises temporal visual and F/T representations by predicting proprioceptive state transitions and is robust to perturbations that substantially degrade pose-based and learned-fusion baselines.

Abstract

Contact-rich assembly remains challenging because it requires submillimeter spatial accuracy and reliable interpretation of forces during sustained contact. Although simulation-based reinforcement learning offers a scalable training paradigm, discrepancies in visual observations, contact dynamics, and force/torque (F/T) measurements often limit policy transfer. We observe that proprioception is comparatively consistent across domains because joint positions are expressed in a shared calibrated coordinate system and joint velocities are computed consistently in simulation and on hardware. Based on this observation, we present PACE (Proprioception-Anchored Cross-Modal Encoder), which supervises temporal visual and F/T representations by predicting proprioceptive state transitions. Static domain-specific factors, including lighting, texture, and sensor bias, contain little information about joint motion; optimizing the proposed objective therefore suppresses their influence on the learned representation while retaining task-relevant motion cues. Policies trained on frozen PACE features are directly deployed on hardware without real-world fine-tuning or object-pose tracking. Across four contact-rich assembly tasks, PACE attains an average real-world success rate of 93.3\% and only a 2.7-percentage-point sim-to-real drop, while remaining robust to perturbations that substantially degrade pose-based and learned-fusion baselines.

View source

Similar papers

Preprint Aug 2026

TRACE: Learned Proprioceptive Odometry for Legged Robots under Unreliable Contact Conditions

In this paper, we present TRACE (Tokenized Robust Attention for Contact-Aware Estimation), an end-to-end learned proprioceptive odometry estimator for legged robots under unreliable contact conditions. The proposed estimator directly predicts relative displacement, relative rotation, and body-frame velocity from a rece...

Taehyeon Kong, W. Kim, Jemin Hwangbo · 0 citations
Preprint Aug 2026

Task-Relevant Feature-Dynamics Fidelity Enables Zero-Shot Sim-to-Real Transfer for Robotic Ultrasound Scanning

Robotic ultrasound policies operating directly on B-mode images require extensive interaction data, whereas real-robot data collection is costly and safety-constrained. Simulation provides a scalable alternative, but zero-shot transfer depends not only on single-frame realism but also on whether simulated observations...

Yi-Zhao Qian, Jia-Yuan Luo, Wang Zhu et al. · 0 citations
Open access 2026

Cross-Modal Distillation With Latent Consistency for Proprioceptive Quadruped Locomotion

Proprioception-only locomotion over complex terrain is essential for quadruped robots when external perception is unavailable or unreliable. Existing methods either rely on privileged physical quantities for supervision or infer environmental information from short-term state evolution, which may lead to overemphasis o...

Kaicen Li, Lingyu Yin, Mengyang Li et al. · 0 citations
Preprint Sep 2026

DAVIS: A Depth-Only End-to-End Active-Vision Framework for Humanoid Soccer Skills

Humanoid soccer contact skills require more than producing high-impact foot-ball contacts: the robot must close the loop over perception, approach, alignment, impact, and recovery while its own motion induces substantial viewpoint changes, frequent loss of the ball from view, and uncertain contact outcomes. In this wor...

Jia-Kang Jin, Yi-Xiao Huo, Peng-Yuan Wang et al. · 0 citations
Preprint Sep 2026

TAO-Force: Unifying Force-Aware Perception and Fast-Slow Control for Contact-Rich Manipulation

Vision-Language-Action (VLA) models have demonstrated strong performance across diverse robotic manipulation tasks, yet their predominantly vision-centric perception and position-controlled execution remain insufficient for contact-rich manipulation. Visual observations alone often provide limited evidence of contact o...

Bo-Han Gan, Xuan Wen, Yong-Shen Zhao et al. · 0 citations
Preprint Sep 2026

Towards High-DoF Dexterous Manipulation through VLA Post-Training

Imitation-learned vision--language--action (VLA) foundation models acquire broad manipulation capabilities by scaling robot data across tasks and embodiments, but reliable deployment on a specific downstream task and hardware platform still requires post-training. Dexterous hands make this adaptation particularly diffi...

Jun-Lei Zhu, Shen-Zhe Yao, Chao-Gui Huang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.