Skip to content
Preprint

AnyViewDex: View-Invariant Dexterous Manipulation from RGB Observations

Sep 2026 · 0 citations · 30 references
Computer Science

TL;DR

This work shows that view-invariant control can be achieved without explicit test-time 3D sensing by encoding geometric knowledge into the visual representation during simulation, and presents AnyViewDex, an asymmetric training pipeline that combines multi-view contrastive alignment with privileged 3D geometric supervision.

Abstract

Visuomotor policies for multi-fingered dexterous manipulation are highly sensitive to camera viewpoint shifts. To achieve view invariance, recent methods increasingly rely on explicit 3D modalities like RGB-D or point clouds, which can introduce hardware dependencies, calibration requirements, and vulnerability to sensor noise during real-world deployment. In this work, we show that view-invariant control can be achieved without explicit test-time 3D sensing by encoding geometric knowledge into the visual representation during simulation. We present AnyViewDex, an asymmetric training pipeline that combines multi-view contrastive alignment with privileged 3D geometric supervision. By regressing absolute 3D object coordinates during simulated training, this auxiliary objective provides a geometric grounding signal that mitigates the spatial collapse of the globally pooled contrastive embedding. At deployment, the policy operates zero-shot using only uncalibrated monocular RGB and proprioception. We validate this approach across both reinforcement learning and student-teacher distillation. In hardware evaluation on an xArm7 with a 16-DoF LEAP Hand, AnyViewDex reaches 76.7% grasping success across eight unseen objects and six uncalibrated viewpoints (480 trials; 2,400 across all ablation conditions), indicating that geometrically grounded monocular policies transfer zero-shot without test-time depth. Project Page: https://anyviewdex.github.io/

View source

Similar papers

Preprint Sep 2026

DEXTERA: From a Single Image to Deployable Dexterous Manipulation via Real-to-Sim-to-Real

Collecting real-world robot data for dexterous manipulation is costly and time-consuming. While high-fidelity physics simulators enable scalable data synthesis and policy learning, constructing deployment-ready digital twins manually remains labor-intensive, and residual visual, geometric, and dynamics gaps hinder reli...

Jin Wu, Lian-Jie Yuan, Ze-Yang Sun et al. · 0 citations
Preprint Sep 2026

StereoPatch: Patch-Aligned RGB-Depth Fusion for Spatial Perception in Robot Manipulation

StereoPatch is introduced, a patch-aligned RGB-depth representation that binds registered metric geometry directly to the RGB patches used for action prediction and suggests that resolving control-relevant geometric ambiguity benefits from aligning depth directly with the visual features used for action prediction.

Ya-Nan Zhou, Zhao-Yan Qian, James Zhao et al. · 0 citations
Preprint Sep 2026

LIFD: Anchored Diffusion for 3D-Aware Scene Memory in Robotic Manipulation

This work introduces \lifd{} (Look, Imagine, Focus, and Do), a framework for persistent, 3D-aware scene memory that learns scene tokens through multi-view agreement, then completes them from a single RGB view and recurrent memory using rectified flow.

Wen-Bo Li, Yi-Teng Chen, Wen-Hao Li et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Emergent Multi-View Geometry Through Self-Distillation

Over a century ago, Henri Poincar\'e argued that a motionless observer cannot acquire the notion of space. Yet, most visual representation learning methods operate on individual images, while those that leverage multiple views rely on RGB reconstruction, entangling geometry with appearance. We propose Poincar3, a self-...

David Nordström, T. Loiseau, Vincent Lepetit et al. · 0 citations
Preprint Aug 2026

ARGUS: Aligning Robot Scene Geometry Under Shifting Views with Large 3D Vision Models

Large-scale visuomotor policies have demonstrated impressive performance across a wide range of robot manipulation tasks. However, despite this success, manipulation polices often entangle scene geometry with the corresponding viewpoint, learning where objects lie in an image rather than where it lies in the task space...

Rishik Sathua, Hao-Nan Chen, K. Driggs-Campbell · 0 citations
Preprint Sep 2026

SyncWorld: Visual Calibration Enables World Models as Zero-Shot Simulators

World models are increasingly used as policy-in-the-loop imagination environments, where reliable rollouts require fine-grained controllability with respect to low-level robot actions. A key obstacle to scaling such models in robotics is that actions are not a universal language in pixel space: changes in visual enviro...

Yuncong Yang, Zhen Han, Furkan Ozyurt et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.