This work shows that view-invariant control can be achieved without explicit test-time 3D sensing by encoding geometric knowledge into the visual representation during simulation, and presents AnyViewDex, an asymmetric training pipeline that combines multi-view contrastive alignment with privileged 3D geometric supervision.
Abstract
Visuomotor policies for multi-fingered dexterous manipulation are highly sensitive to camera viewpoint shifts. To achieve view invariance, recent methods increasingly rely on explicit 3D modalities like RGB-D or point clouds, which can introduce hardware dependencies, calibration requirements, and vulnerability to sensor noise during real-world deployment. In this work, we show that view-invariant control can be achieved without explicit test-time 3D sensing by encoding geometric knowledge into the visual representation during simulation. We present AnyViewDex, an asymmetric training pipeline that combines multi-view contrastive alignment with privileged 3D geometric supervision. By regressing absolute 3D object coordinates during simulated training, this auxiliary objective provides a geometric grounding signal that mitigates the spatial collapse of the globally pooled contrastive embedding. At deployment, the policy operates zero-shot using only uncalibrated monocular RGB and proprioception. We validate this approach across both reinforcement learning and student-teacher distillation. In hardware evaluation on an xArm7 with a 16-DoF LEAP Hand, AnyViewDex reaches 76.7% grasping success across eight unseen objects and six uncalibrated viewpoints (480 trials; 2,400 across all ablation conditions), indicating that geometrically grounded monocular policies transfer zero-shot without test-time depth. Project Page: https://anyviewdex.github.io/
Collecting real-world robot data for dexterous manipulation is costly and time-consuming. While high-fidelity physics simulators enable scalable data synthesis and policy learning, constructing deployment-ready digital twins manually remains labor-intensive, and residual visual, geometric, and dynamics gaps hinder reli...
Jin Wu, Lian-Jie Yuan, Ze-Yang Sun et al.· 0 citations
StereoPatch is introduced, a patch-aligned RGB-depth representation that binds registered metric geometry directly to the RGB patches used for action prediction and suggests that resolving control-relevant geometric ambiguity benefits from aligning depth directly with the visual features used for action prediction.
Ya-Nan Zhou, Zhao-Yan Qian, James Zhao et al.· 0 citations
This work introduces \lifd{} (Look, Imagine, Focus, and Do), a framework for persistent, 3D-aware scene memory that learns scene tokens through multi-view agreement, then completes them from a single RGB view and recurrent memory using rectified flow.
Wen-Bo Li, Yi-Teng Chen, Wen-Hao Li et al.· 0 citations
Over a century ago, Henri Poincar\'e argued that a motionless observer cannot acquire the notion of space. Yet, most visual representation learning methods operate on individual images, while those that leverage multiple views rely on RGB reconstruction, entangling geometry with appearance. We propose Poincar3, a self-...
David Nordström, T. Loiseau, Vincent Lepetit et al.· 0 citations
Large-scale visuomotor policies have demonstrated impressive performance across a wide range of robot manipulation tasks. However, despite this success, manipulation polices often entangle scene geometry with the corresponding viewpoint, learning where objects lie in an image rather than where it lies in the task space...
Rishik Sathua, Hao-Nan Chen, K. Driggs-Campbell· 0 citations
World models are increasingly used as policy-in-the-loop imagination environments, where reliable rollouts require fine-grained controllability with respect to low-level robot actions. A key obstacle to scaling such models in robotics is that actions are not a universal language in pixel space: changes in visual enviro...
Yuncong Yang, Zhen Han, Furkan Ozyurt et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.