Skip to content
Preprint

Zero-WAM: In-Context World-Action Modeling from Human Videos for Open-Ended Task Generalization

Aug 2026 · 8 citations · 84 references
Computer Science

TL;DR

This work presents Zero-WAM, a causal video-action model that executes unseen tasks by following in-context human video guidance, and proposes an automatic pipeline that converts task-sampled robot trajectories into semantically matched human videos, yielding HumanGen, a dataset of 74.2K human-robot ICL pairs across 8.6K tasks.

Abstract

Zero-shot cross-task generalization, where a policy must execute manipulation tasks never seen during training, remains a central challenge in robot learning. In large language models, a novel task can be performed simply by specifying it in the context, without any parameter update. This form of in-context learning (ICL) turns generalization into a problem of task specification. To achieve cross-task generalization, we bring this paradigm to robotic manipulation, and argue that the natural task specification for manipulation is a human video: unlike language, it provides rich visual cues about the intended task evolution. We present Zero-WAM, a causal video-action model that executes unseen tasks by following in-context human video guidance. To address the scarcity of task-rich paired human-robot data, we propose an automatic pipeline that converts task-sampled robot trajectories into semantically matched human videos, yielding HumanGen, a dataset of 74.2K human-robot ICL pairs across 8.6K tasks. For model training, we further introduce an in-context future chunk prediction (IFP) objective that suppresses shortcuts learned from seen tasks and forces the policy to draw task information from the video prompt. On seven unseen tasks in RoboTwin 2.0 simulation, Zero-WAM achieves a 47.0% average success rate, an absolute improvement of 29.5 percentage points over the strongest video-action baseline. In real-world evaluations, it follows human video guidance to generalize to unseen task configurations involving multi-object scenes, long-horizon manipulation, and fine-grained insertion.

View source

Similar papers

Preprint Sep 2026

In-context Robot Learning Made Simple: A Democratized Recipe for Manipulation Tasks

We study robotic in-context learning (ICL), an emerging paradigm that enables robots to infer and execute tasks from visual demonstrations. Despite its growing promise, the problem itself remains under-defined: a visual demonstration simultaneously conveys action trajectories, object semantics, manipulation affordances...

Min-Xing Li, Ming-Hao Han, Wei-Zhi Zhao et al. · 0 citations
Preprint Sep 2026

MotorMind: Scaffolding General Vision Language Models for Zero-Shot Robot Manipulation

Vision-language-action (VLA) models have advanced robotic manipulation, but their zero-shot generalization in new tasks and environments remains limited, and their reliance on specialized training keeps them from benefiting directly from rapidly advancing general-purpose vision-language models (VLMs). In parallel, rece...

Bing-Xuan Li, Si-Qi Song, Yi-Zhuo Wu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

HIL-UMI: Bringing Human-in-the-Loop Post-Training of Vision-Language-Action Models to Universal Manipulation Interface

Large-scale vision-language-action (VLA) models provide powerful priors for robot manipulation, yet adapting them to a specific deployment remains challenging. Supervised fine-tuning (SFT) on task-specific demonstrations provides a step toward deployment, but faces two persistent limitations: static data provide limite...

Zimu Han, Yi-Ming Zeng, Ji-Yao Zhang et al. · 0 citations
Preprint Sep 2026

EVO-WAM: Evolving World Action Models through Video-Action Verification

Improving robot policies on new tasks without collecting additional expert demonstrations remains a central challenge in robot learning. World action models (WAMs) use broad video priors to jointly predict future videos and actions, offering a potential source of supervision for adapting to new tasks. However, generate...

Shi-Long Zhou, Xiong-Hao Wu, Wen-Bo Li et al. · 0 citations
Preprint Aug 2026

JoyAI-RA 0.5: Scaling Robot Manipulation Learning via Dual Action Alignment

JoyAI-RA 0.5 is proposed, a generalist Vision-Language-World-Action framework that couples physical world-dynamics priors with visual semantics and scales manipulation learning across such data via dual action alignment, suggesting that abundant but weakly labeled human experience can be converted into a transferable t...

JoyAI-RA Team · 3 citations · ⚡1

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.