Skip to content
Preprint

AR-WAM: A Visual-Conditioned Agent-Ready World Action Model for Robotic Manipulation

Sep 2026 · 0 citations · 51 references
Computer Science

TL;DR

AR-WAM is presented, a visual-conditioned, agent-ready world action model that replaces language with two complementary conditions: a visual grounding prompt and a learnable operation token dictating the atomic skill to execute, predicting scene evolution within compact latent states while decoding actions.

Abstract

As AI agents become increasingly capable, agent-driven robotic control is emerging as a compelling paradigm. However, prevailing vision-language-action (VLA) models and world action models (WAMs) still rely on natural-language instructions to specify manipulation tasks, an ill-suited interface for agent-driven control: referentially ambiguous, spatially imprecise, redundant with the agent's inherent language understanding, and entangling intent with execution. We present AR-WAM, a visual-conditioned, agent-ready world action model that replaces language with two complementary conditions: a visual grounding prompt (a bounding box of the target) denoting the interaction object and location, and a learnable operation token dictating the atomic skill to execute. Our compact 0.5B-parameter model, with a frozen pretrained visual encoder and no language encoder, predicts scene evolution within compact latent states while decoding actions, exposing the policy's intent through explicit, supervisable reasoning signals. A model-agnostic compatibility layer provides three primitives (detect, execute, and query) so that local VLMs or online agent APIs can drive the policy directly, with long-horizon memory and closed-loop error recovery delegated to the agent side. On RoboTwin 2.0, RMBench, and a real Astribot S1 dual-arm platform, AR-WAM matches the strongest baselines on standard manipulation (87.2% average success) and outperforms them on memory-dependent and real-robot long-horizon tasks, improving success rates by 5.9% and 36.7%, respectively, while maintaining the lowest inference latency (14.1 ms).

View source

Similar papers

Preprint Sep 2026

SAVLA: Symmetry-Aware Vision-Language-Action Models for Robotic Manipulation

Vision-language-action (VLA) models have become the dominant paradigm for language-conditioned robot manipulation. However, although images and language instructions inherently encode geometric information, VLAs acquire their spatial competence purely from demonstrations. As a result, they are reliable only within the...

Jun-Le Li, Weixian Waylon Li, Fu-Xiang Wu et al. · 0 citations
Preprint Sep 2026

HINT: Human-Intent Inception for Long-Horizon Robot Manipulation

Humans can perform complex manipulations given a simple intent through an overall instruction, while continuously adapting to evolving visual observations. However, current vision-language action (VLA) models and other action policies struggle to realize this high-level intelligent behavior under dense, evolving visual...

Ming-Yu Mei, Haojie Xu, Shi-Hao Jin et al. · 0 citations
#artificial intelligence Review Sep 2026

World Action Agent: Harnessing VLMs for Robot Manipulation via World Action Rehearsal

World Action Agent is presented, a multi-agent harness through which VLMs pilot robots with basic tools, making every decision within a visual action workspace, outperforming end-to-end VLAs, code-as-policy agents, and a visual-harness baseline with the same backbone.

Ye-Hang Zhang, Hao-Jian Huang, Yi-Fan Chang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

TAO-DA: Towards Autonomous Operation--A Dual-Arm Vision-Language-Action Model for Coordinated Manipulation

A symmetric Dual-Arm Expert (DAE) architecture built upon a shared Vision-Language Model (VLM) backbone with decoupled, arm-specific expert towers is proposed, providing preliminary evidence of emergent skill generalization from single- to dual-arm tasks (as well as the reverse), together with cross-arm motion-domain s...

Yong-Shen Zhao, Han Gao, Bao-Ping Cheng et al. · 0 citations
Preprint Aug 2026

LiLa-WAM: Lightweight Latent Reasoning World-Action Model for Robotic Manipulation

LiLa-WAM is proposed, a lightweight world-action model that reasons about the future in a compact latent space and can be trained end-to-end on a single 24GB GPU and the Visual Transition Token (VTT), a language-free task representation that encodes each task as a direction in visual feature space.

Fan Yang, Yu-Ting Su, Xiaobo Wang et al. · 9 citations
Open access Aug 2026

PFEA: a VLM-based high-level natural language planning and feedback embodied agent for human-centered AI

A closed-loop framework for planning and evaluation of a vision-language model-based robotic manipulation agent operating in tabletop object rearrangement and manipulation tasks and demonstrates the potential of closed-loop vision-language planning for human-centered robotic manipulation.

Wenbin Ding, Jun Chen, Mingjia Chen et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.