Zeva-Ego is introduced, a unified framework that learns physical priors from human experience and evolves through robot interaction and demonstrates a scalable path toward embodied intelligence that learns from human experience and continuously improves through its own interaction.
Abstract
Egocentric video offers a scalable source of physical interaction experience, yet translating it into robot-executable knowledge and enabling continual adaptation remain challenging. We introduce Zeva-Ego, a unified framework that learns physical priors from human experience and evolves through robot interaction. An Action-Centric Encoder (ACE) converts egocentric visual transitions into action-centered supervision for VLA mid-training, while In-Context Causal Learning (ICCL) enables parameter-free adaptation from action-effect feedback at deployment. Scaling Ego data to 10K hours improves RoboTwin success from 63.8% to 75.3%, matching 2K hours of robot demonstrations (74.7%), corresponding to an empirical data ratio of roughly 4-5:1. With accumulated interaction experience, ICCL further improves success from 58% to 89% within four attempts without parameter updates. These results demonstrate a scalable path toward embodied intelligence that learns from human experience and continuously improves through its own interaction.
JoyAI-RA 0.5 is proposed, a generalist Vision-Language-World-Action framework that couples physical world-dynamics priors with visual semantics and scales manipulation learning across such data via dual action alignment, suggesting that abundant but weakly labeled human experience can be converted into a transferable t...
This work introduces ZimaBlue, a scalable framework for learning generalizable World Action Models (WAMs) from large-scale video, and adopts an asynchronous Slow-Fast dual-system architecture to make generative WAMs practical for real-time control.
Xiong-Hao Wu, Yi-Jun Yang, Shi-Long Zhou et al.· 2 citations
AnyWorld is proposed, a cross-embodiment world modeling framework that expands a single human interaction into diverse robot-native rollouts without paired human-robot demonstrations and enables independent recomposition of embodiment, viewpoint, and scene factors, allowing a single model to generate many robot-domain...
Cheng Chen, J. Bai, Jiacheng Wei et al.· 3 citations
AtomEgo is presented, a systematic study of ego--robot co-training supported by a curated corpus of approximately 2,659 hours and a scalable data processing pipeline that reveals a simple principle: Data Scale * Alignment Quality -->Capability Gain; egocentric data can improve generalization, but their value depends on...
Di Wu, Dong-Chen Zheng, Jun-He Sheng et al.· 0 citations
Learning generalizable robot manipulation policies requires large-scale and diverse demonstration data. Egocentric human manipulation videos offer rich scene and task diversity, and prior work has shown that retargeting and rendering such videos into robot-format data can yield effective per-task policies at small scal...
Ye Wang, Peibin Lin, Xiong-Hui Chen et al.· 10 citations· ⚡1
Vision--language--action (VLA) models provide strong priors for robotic manipulation but are typically deployed as frozen policies, unable to improve from their own failures. Real-world reinforcement learning (RL) offers a path to continued improvement, yet manual environment resets and task-success supervision hinder...
Yuan Fang, Ze-Chu Li, Hao-Lei Tong et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.