Skip to content
Preprint

Multi-View Unified Camera Fields: Geometry-Shaped Action-Facing Representations for RGB-Only Multi-Camera VLA Policies

Aug 2026 · 1 citation · 28 references
Computer Science

TL;DR

Multi-View Unified Camera Fields (MVUCF), a training-only framework that forms a shared action-facing latent field across views, and real-world humanoid experiments further provide evidence of its practical effectiveness under RGB-only deployment.

Abstract

Vision-Language-Action (VLA) models have shown strong generalization in robotic manipulation, yet complex contact-rich tasks often benefit from multi-camera observations that jointly capture the end effector, objects, and targets under occlusion. Existing multi-camera VLAs usually concatenate view tokens, leaving action representations weak in metric depth and inconsistent across cameras. We introduce Multi-View Unified Camera Fields (MVUCF), a training-only framework that forms a shared action-facing latent field across views. A coordinate-query depth objective makes metric depth recoverable, while a preprocessing-aware correspondence objective aligns tokens observing the same physical point from different cameras. Both directly shape the hidden states consumed by the action module. After geometry injection, depth, camera calibration, and auxiliary heads are removed, so deployment uses the original RGB-only graph with no extra inference FLOPs. Held-out probes confirm stronger depth recovery and cross-view matching. Under matched GR00T-N1.6 settings, MVUCF reaches 98.9% on LIBERO, improves LIBERO-Plus by 22.4 points, and raises success by 23.3 points across six RoboTwin tasks spanning three action families: touch, move-and-place, and contact interaction. Real-world humanoid experiments further provide evidence of its practical effectiveness under RGB-only deployment.

View source

Similar papers

Preprint Aug 2026

Cross-View Action Consistency for Camera-Robust Vision-Language-Action Policies

Vision-language-action (VLA) policies fine-tuned from a fixed scene camera can fail when the camera is moved, even when the task, objects, language, and robot state are unchanged. We study scene-camera viewpoint robustness using only a scene RGB image, language, and proprioception, without camera labels, extrinsics, de...

Bingqi Huang, Bingchuan Wei, Xuan Wang et al. · 3 citations
Preprint Sep 2026

MVG-WAM: Multiple View Geometry-Aware World-Action Modeling for Robotic Manipulation

World-Action Models (WAMs) couple visual dynamics with action prediction, bringing the rich priors of pretrained video models to robotic manipulation. However, their multi-view interfaces typically tile images or concatenate tokens, leaving the geometric relationships among synchronized cameras implicit. This makes it...

Wen-Bo Chen, Tian-Fu Li, Hao-Xuan Xu et al. · 0 citations
Preprint Sep 2026

M3GD: Multi-Modal Multi-View Geometric Diffusion for Camera--LiDAR Novel View Synthesis

M3GD is presented, a Camera--LiDAR multimodal representation for generative NVS that composes independently pretrained 2D image and 3D point-cloud foundation models without separately pretraining a cross-modal translator.

Yang Zhou, Jiuhong Xiao, Shi-Zhao Ye et al. · 0 citations
Preprint Aug 2026

Selective Cross-View Consistency for World Action Models: Held-Out Viewpoint Robustness Without Test-Time Camera Information

World action models (WAMs) jointly denoise future video frames and robot actions, and the video prior is expected to generalize their control. Camera viewpoint change remains one of their hardest perturbation axes. We study a question specific to this model class: when training with same-state cross-view image pairs, o...

Bingqi Huang, Bingchuan Wei, Ying-Kai Cai et al. · 2 citations
Preprint Sep 2026

MINT: A Unified Model for World-Space Camera and Hand Motion Estimation from Scalable Egocentric Pipeline Supervision

This work introduces MINT (Minting IN-the-Wild Trajectories), a foundation model for world-space hand motion reconstruction from ego-centric RGB video and develops an open-source labeling EGOPIPELINE that converts large collections of public egocentric videos into structured camera-and-hand trajectory supervision.

Zi-Jie Zhu, Wei-Ren Cai, Yi-Zhou Wang et al. · 3 citations
Preprint Sep 2026

LIFD: Anchored Diffusion for 3D-Aware Scene Memory in Robotic Manipulation

This work introduces \lifd{} (Look, Imagine, Focus, and Do), a framework for persistent, 3D-aware scene memory that learns scene tokens through multi-view agreement, then completes them from a single RGB view and recurrent memory using rectified flow.

Wen-Bo Li, Yi-Teng Chen, Wen-Hao Li et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.