Skip to content
Preprint

EgoWild2Dex: Learning Dexterous Robotic Manipulation from In-the-Wild Human Experience

Sep 2026 · 0 citations
Computer Science

TL;DR

EgoWild2Dex, which transfers in-the-wild ego-human experience to dual-arm robots with dexterous hands by jointly aligning unstable egocentric views and human motions with robot observations and actions, is introduced and GeoFormer, a differentiable geometric transformer that warps noisy human observations toward robot observations, is introduced.

Abstract

Egocentric human data provide a principled source of supervision for learning dexterous robot manipulation. Unlike prior approaches that often collect such data in constrained or specially constructed environments, we collect in-the-wild egocentric demonstrations in real-world settings, including homes, factories, and pharmacies, etc., where people perform their ordinary tasks while wearing head-mounted cameras. This collection protocol captures diverse workflows and hand-object interactions across long-tailed object and skill distributions, but also yields visually challenging observations due to scene clutter and head-motion-induced viewpoint changes (a mean cumulative rotation of $15.93^{\circ}$/s). To address these issues, we introduce EgoWild2Dex, which transfers in-the-wild ego-human experience to dual-arm robots with dexterous hands by jointly aligning unstable egocentric views and human motions with robot observations and actions, respectively. This work offers three benefits. First, we introduce GeoFormer, a differentiable geometric transformer that warps noisy human observations toward robot observations. Second, we design a human-robot training scheme to bridge the embodiment gap, enabling high task success with limited robot supervision. Third, we release EgoWild, a 538.9-hour in-the-wild egocentric human dataset comprising 179,049 episodes, 125,961 unique task descriptions, and 1,282 object categories. On real robots, EgoWild2Dex achieves an average success rate of 96.7% across three long-horizon bimanual dexterous manipulation tasks and an average object-level zero-shot success rate of 33.3%. The data, models, and code will be released.

View source

Similar papers

Preprint Aug 2026

HandEdit: A Unified Benchmark for Egocentric Human-to-Robot Dexterous Hand Image Editing

Robotic manipulation with dexterous hands is a cornerstone of Embodied AI, yet its progress is stifled by the high cost of collecting embodiment-aware teleoperation data. While abundant egocentric videos of human hands offer a scalable alternative, the profound discrepancies in appearance, articulation, and camera view...

Zhen-Jie Yang, Xingyu Jiao, Guopeng Zhong et al. · 5 citations
Preprint Aug 2026

Ego2Robot: Scalable Robot Data Synthesis from Egocentric Human Data

Learning generalizable robot manipulation policies requires large-scale and diverse demonstration data. Egocentric human manipulation videos offer rich scene and task diversity, and prior work has shown that retargeting and rendering such videos into robot-format data can yield effective per-task policies at small scal...

Ye Wang, Peibin Lin, Xiong-Hui Chen et al. · 10 citations · ⚡1
Preprint Aug 2026

RoboReact: Agentic Skill Distillation from Generated Egocentric Videos for Generalizable Whole-Body Manipulation

This work presents RoboReact, a framework that automatically synthesizes whole-body humanoid manipulation skills from a single egocentric RGB-D observation, and highlights the potential of combining generative models, vision-language reasoning, and closed-loop control for scalable humanoid skill acquisition.

Shu-Liang He, Shuai Wang, Bo Yue et al. · 3 citations · ⚡1
Preprint Aug 2026

EgoTac: In-the-wild Tactile Prediction from Egocentric Vision

Touch is fundamental to dexterous manipulation, yet most egocentric human data increasingly used for robot learning lacks tactile information. Directly collecting large-scale tactile data is challenging due to sensor limitations, while human video data is abundant, contact-rich, and easily scalable. This motivates a na...

Wen-Kang Zhang, Chengbo Yuan, Zicheng Zhang et al. · 1 citation · ⚡1
Preprint Sep 2026

DexRoam: Learning Mobile Bimanual Dexterous Manipulation from Egocentric Whole-Body Human Demonstrations

Mobile bimanual dexterous manipulation requires continuous coordination of locomotion, whole-body motion, and finger-level dexterity within a single trajectory, creating a severe robot demonstration bottleneck. Egocentric human demonstrations offer a scalable alternative, but prior approaches ease the transfer by simpl...

Rui Zhou, Yi-Bo Yuan, Jun-Kai Zhao et al. · 0 citations
Feb 2025

RobotMover: Learning to Move Large Objects From Human Demonstrations

Moving large objects, such as furniture or appliances, is a critical capability for robots operating in human environments. This task presents unique challenges, including whole-body coordination to avoid collisions and managing the underactuated dynamics of bulky, heavy objects. In this work, we present RobotMover, a...

Tian-Yu Li, Joanne Truong, Tsung-Yen Yang et al. · 2 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.