Skip to content
Preprint

MINT: A Unified Model for World-Space Camera and Hand Motion Estimation from Scalable Egocentric Pipeline Supervision

Sep 2026 · 3 citations · 40 references
Computer Science

TL;DR

This work introduces MINT (Minting IN-the-Wild Trajectories), a foundation model for world-space hand motion reconstruction from ego-centric RGB video and develops an open-source labeling EGOPIPELINE that converts large collections of public egocentric videos into structured camera-and-hand trajectory supervision.

Abstract

Recovering camera and hand motion in world coordinates from egocentric video is a key capability for activity understanding, robot learning, and augmented reality. Existing systems typically decompose this problem into separate stages for camera motion, depth estimation, hand reconstruction, and trajectory refinement, resulting in substantial computational overhead and preventing the joint modeling of camera and hand motion. We introduce MINT (Minting IN-the-Wild Trajectories), a foundation model for world-space hand motion reconstruction from ego-centric RGB video. From a single shared spatiotemporal video representation, MINT jointly predicts the camera trajectory, field of view (FoV), camera-frame hand states, and per-frame hand observability, and then produces world-space hand motion via explicit coordinate transformations. Training such a model at scale is challenging, since paired world-space camera and hand annotations are scarce. We therefore develop an open-source labeling EGOPIPELINE that converts large collections of public egocentric videos into structured camera-and-hand trajectory supervision. MINT is first pretrained on these large-scale pseudo-labels and then fine-tuned on a small set of high-quality camera-and-hand annotations. Across public benchmarks MINT approaches state-of-the-art accuracy without seeing either benchmark in training, reaching 0.945 frame accuracy, 13.646 mm PA-MPJPE-p and 55.058 px EPE-p for camera-frame bimanual reconstruction on HOT3D, 4.690 mm RPE-T and 0.284 degrees RPE-R for camera trajectory, and a 3.67x end-to-end speedup over the labeling pipeline that supervises it. We release the model, training and inference code, labeling pipeline, and a curated 1,021-hour egocentric trajectory dataset.

View source

Similar papers

Preprint Sep 2026

InfiniHand: Streaming World-Space Hand Motion Estimation from Egocentric Video

World-space hand motion estimation from egocentric video requires recovering 3D articulated hand geometry while tracking camera egomotion. Existing approaches heavily rely on cascading independent hand pose estimators and SLAM systems, resulting in error accumulation, complex pipelines, and severe computational overhea...

Ke-Rui Ren, Kai-Wen Song, Weiguang Zhao et al. · 0 citations
Preprint Aug 2026

ACE-Ego-Hand: Repurposing Video Diffusion Models for Occlusion-Robust Egocentric 3D Hand Motion Recovery

ACE-Ego-Hand is introduced, an offline clip-level framework that extracts features via a Deterministic Clean-Latent Encoder and decodes them with a Bidirectional Spatiotemporal Decoder, offering a scalable path from everyday human video to robot manipulation data.

Yu-Fei Liu, Xixi Wang, Hao Li et al. · 1 citation
Preprint Sep 2026

Seeing the World and the Self from Egocentric Video

Complete 3D perception from egocentric video requires recovering the surrounding scene and the wearer's full-body motion in a shared metric frame. Existing methods typically address scene reconstruction and motion estimation separately: scene reconstruction methods ignore the wearer, whereas motion estimation methods l...

Kai Guan, Minchao Jiang, Ruichen WangLi et al. · 0 citations
#artificial intelligence Preprint Sep 2026

DiffWAM: A Fast and Efficient Navigation World Action Model

Pretrained video foundation models encode rich semantic and spatiotemporal priors for embodied navigation, yet converting these priors into UAV motion typically requires expensive future-video synthesis and geometric reconstruction. We investigate whether the motion implicit in future visual prediction can instead be r...

Morui Zhu, Yu-Ze Wu, Xi-Jie Huang et al. · 0 citations
2026

GCAFormer-VO: A Geometry and Correspondence-Aware Transformer for Visual Odometry

Visual odometry (VO) estimates camera motion from image sequences and is essential for robotics, autonomous driving, and AR/VR. Robust VO remains challenging because large viewpoint changes and strong parallax make reliable cross-frame motion cues difficult to capture, especially in the presence of visual disturbances...

Jun-Qi Bao, Qing-Ying Wu, Jun Huang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Estimating Accurate Hand Pose in Camera Space with Vision Transformer

Monocular RGB-based hand pose estimation has emerged as a critical research frontier in computer vision. The local hand pose estimation methods predict hand poses relative to the wrist, while global hand pose estimation also requires estimating the wrist's position in the camera coordinate system. However, this camera-...

Kai-Wen Ren, Yi-Ran Jiang, Yong-Jing Ye et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.