Skip to content

World Action Agent: Harnessing VLMs for Robot Manipulation via World Action Rehearsal

Sep 2026 · 0 citations · 39 references
Computer Science

TL;DR

World Action Agent is presented, a multi-agent harness through which VLMs pilot robots with basic tools, making every decision within a visual action workspace, outperforming end-to-end VLAs, code-as-policy agents, and a visual-harness baseline with the same backbone.

Abstract

General-purpose vision-language models (VLMs) bring broad knowledge and spatial reasoning to robot manipulation, yet existing systems either use them indirectly, to predict constraints or write programs, or give them a view of the scene rather than a world in which to act. We present World Action Agent (WAA), a multi-agent harness through which VLMs pilot robots with basic tools, making every decision within a visual action workspace. The workspace has three properties. Contact views, selected automatically from the scene geometry, present the scene around the current interaction. Action rehearsal turns each action into an editable proposal that the agent, alone or through an Imagination Agent, previews and revises against planning feedback before execution. In-view correction closes the loop between observation, rehearsal, and low-level execution, letting the agent remove residual offsets in the view where it observes them. Through the same workspace, WAA acquires embodied procedural knowledge in two ways: it evolves multimodal skills from expert videos and human teaching under evidence-based review and consults them through a Skill Agent, and its interaction traces train smaller VLMs to pilot the same harness. On LIBERO-Pro, WAA with skills evolved only from LIBERO-90 reaches a state-of-the-art 75.6% average success, outperforming end-to-end VLAs, code-as-policy agents, and a visual-harness baseline with the same backbone; the same skills remain effective on robosuite without further learning. Fine-tuning Qwen3.5-9B on harness traces raises its out-of-domain success from 1.7% to 43.3%.

View source

Similar papers

Preprint Sep 2026

AR-WAM: A Visual-Conditioned Agent-Ready World Action Model for Robotic Manipulation

AR-WAM is presented, a visual-conditioned, agent-ready world action model that replaces language with two complementary conditions: a visual grounding prompt and a learnable operation token dictating the atomic skill to execute, predicting scene evolution within compact latent states while decoding actions.

Yi-Cheng Jiang, Ze-Sen Gan, Xiao-Bo Wang et al. · 0 citations
Preprint Sep 2026

Find Something You Can't Do: Agentic Real-World Reinforcement Learning for Self-Improving VLA Models

Vision--language--action (VLA) models provide strong priors for robotic manipulation but are typically deployed as frozen policies, unable to improve from their own failures. Real-world reinforcement learning (RL) offers a path to continued improvement, yet manual environment resets and task-success supervision hinder...

Yuan Fang, Ze-Chu Li, Hao-Lei Tong et al. · 0 citations
Preprint Sep 2026

HINT: Human-Intent Inception for Long-Horizon Robot Manipulation

Humans can perform complex manipulations given a simple intent through an overall instruction, while continuously adapting to evolving visual observations. However, current vision-language action (VLA) models and other action policies struggle to realize this high-level intelligent behavior under dense, evolving visual...

Ming-Yu Mei, Haojie Xu, Shi-Hao Jin et al. · 0 citations
#artificial intelligence Preprint Sep 2026

2AM: Grounding Agent-Side Memory as Guidance for Steerable Action Models in Long-Horizon Manipulation

Long-horizon robot manipulation requires memory, but not necessarily inside the action policy. To address such tasks, current agentic systems often combine VLAs with planners and geometric tools, sometimes using additional depth or calibrated geometry. These systems confound attribution: gains may come from richer obse...

Yu-Tong Hu, Feng Chen, Xue-Zhi Cao et al. · 1 citation · ⚡1
Preprint Sep 2026

Think Like a World Model, Act Like a VLA: Distilling World-Model Representations into Compact Robot Policies

Vision-Language-Action (VLA) models map observations to actions with no objective that accounts for how the world responds, so their robustness is bounded primarily by data coverage. World models carry precisely that missing objective and are better grounded for it, yet rolling the future forward costs seconds per deci...

T. Dao, Sankalp Yamsani, Jaden Park et al. · 0 citations
Preprint Sep 2026

HarnessVLN: Unifying Training-Free Embodied Navigation through an Agent Harness

Embodied navigation requires agents to ground instructions or object goals in spatial observations and translate plans into successful execution. As multimodal large language models (MLLMs) become increasingly capable, they offer stronger support for navigation without task-specific training; however, improved semantic...

Yang Chen, Li-Rong Che, Zhen-Yu Huang et al. · 4 citations · ⚡1

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.