Skip to content

AtomEgo: Exploring Ego-Robot Integration for Embodied Foundation Model Pretraining

Sep 2026 · 0 citations · 39 references
Computer Science

TL;DR

AtomEgo is presented, a systematic study of ego--robot co-training supported by a curated corpus of approximately 2,659 hours and a scalable data processing pipeline that reveals a simple principle: Data Scale * Alignment Quality -->Capability Gain; egocentric data can improve generalization, but their value depends on how effectively they are aligned and utilized.

Abstract

Embodied foundation models are constrained by the limited scale and diversity of robot demonstrations, motivating the use of large-scale egocentric human interaction data. However, how to effectively incorporate such data into embodied-model pre-training remains unclear because of substantial embodiment and action-space gaps between humans and robots. We present AtomEgo, a systematic study of ego--robot co-training supported by a curated corpus of approximately 2,659 hours and a scalable data processing pipeline. Across vision--language--action and world--action model architectures, we investigate three representative paradigms: joint co-training with domain-specific action heads, progressive ego-to-robot transfer through embodiment alignment, and joint video--action modeling. We evaluate these paradigms through multi-task real-robot experiments and language-conditioned cross-embodiment representation analysis. Our results reveal a simple principle: Data Scale * Alignment Quality -->Capability Gain; egocentric data can improve generalization, but their value depends on how effectively they are aligned and utilized. This principle can provide practical guidance for scalable ego--robot pre-training.

View source

Similar papers

Preprint Aug 2026

AnyWorld: Factorized Egocentric World Models for Cross-Embodiment Generalization

AnyWorld is proposed, a cross-embodiment world modeling framework that expands a single human interaction into diverse robot-native rollouts without paired human-robot demonstrations and enables independent recomposition of embodiment, viewpoint, and scene factors, allowing a single model to generate many robot-domain...

Cheng Chen, J. Bai, Jiacheng Wei et al. · 3 citations
Preprint Aug 2026

Ego2Robot: Scalable Robot Data Synthesis from Egocentric Human Data

Learning generalizable robot manipulation policies requires large-scale and diverse demonstration data. Egocentric human manipulation videos offer rich scene and task diversity, and prior work has shown that retargeting and rendering such videos into robot-format data can yield effective per-task policies at small scal...

Ye Wang, Peibin Lin, Xiong-Hui Chen et al. · 10 citations · ⚡1
Preprint Sep 2026

EgoAlign: Bridging the Human-Humanoid Gap for Long-Range Loco-Manipulation

Egocentric human demonstrations offer an accessible source of task experience, but differences in body scale and controller response, together with missing robot states, limit their value as humanoid training supervision. We present EgoAlign, a data-construction framework that converts these demonstrations into action...

Yi-Ming Jiang, Jin Chen, Chong-Yang Xu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

EgoHumanoid-V2: Human-to-Humanoid Transfer of Coordinated Whole-Body Skills for Loco-Manipulation

Human demonstrations capture diverse scenes and rich whole-body skills without requiring robot teleoperation. Prior work on egocentric transfer has emphasized scene generalization in loco-manipulation under decoupled control, leaving direct transfer of coordinated whole-body skills less explored. We present EgoHumanoid...

Jin Chen, Yi-Ming Jiang, Chong-Yang Xu et al. · 0 citations
Preprint Sep 2026

MachEmbodied-U0: Unified Understanding and Generation Model for Embodied Intelligence

General-purpose robot control requires models to understand task intent, identify where to interact, capture how the scene evolves, and generate precise actions. Vision-language-action models provide strong semantic priors but typically do not explicitly model scene dynamics, while world-action models couple visual pre...

Hao-Ran Wen, Wen-Fu Wang, Kun-Song Shi et al. · 1 citation
Preprint Sep 2026

Zeva-Ego: Egocentric Mid-Training with In-Context Causal Learning for Robot Manipulation

Zeva-Ego is introduced, a unified framework that learns physical priors from human experience and evolves through robot interaction and demonstrates a scalable path toward embodied intelligence that learns from human experience and continuously improves through its own interaction.

Bing-Jia Huang, Xin Ding, Fu Chen et al. · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.