Skip to content

CodeActionBench: Evaluating Agentic Code-as-Policy for Embodied Manipulation

Sep 2026 · 0 citations · 63 references
Computer Science

TL;DR

CodeActionBench is introduced, a benchmark of 25 manipulation tasks that evaluates this capability through agentic Code-as-Policy and provides a controlled testbed for measuring how general-purpose models translate their capabilities into manipulation behavior and for examining typical failure scenarios in that process.

Abstract

How well can general-purpose multimodal models turn visual understanding and reasoning into embodied manipulation via executable code? We introduce CodeActionBench, a benchmark of 25 manipulation tasks that evaluates this capability through agentic Code-as-Policy. Without task-specific fine-tuning, demonstrations, external specialist perception or grasp modules, privileged scene state, or predefined task policies, agents should select visual evidence, form task-relevant 3D estimates, construct manipulation targets, and iteratively execute and revise their policies. A shared robot API provides RGB observations, calibrated geometric operations, robot feedback, and bounded motion, leaving task-dependent decisions to the evaluated agent. Fixed task instances, resource budgets, and a hidden physical-outcome verifier support controlled comparisons across models and harness configurations. Extensive evaluations across nine configurations and 675 attempts achieve success rates ranging from 2.7% to 73.3%. The strongest configuration, GPT-6 Astra with Codex CLI, solves 22 of 25 tasks at least once in three attempts, demonstrating the best performance while still leaving substantial room for improvement. Trajectory analyses reveal difficulties in spatial alignment, object retention, and completion judgment, including task failures despite successfully completed motions. CodeActionBench provides a controlled testbed for measuring how general-purpose models translate their capabilities into manipulation behavior and for examining typical failure scenarios in that process.

View source

Similar papers

Preprint Sep 2026

RACaP: Agentic Reasoning, Acting, and Coding as Policies for Evolvable Robot Learning

General-purpose robot agents must learn from experience, transfer to new tasks, and act efficiently. Code as Policies (CaP) methods generate and repair programs at runtime, incurring latency and entangling reusable mechanisms with task-specific decisions. We introduce RACaP, an agentic framework that moves coding to ev...

Ze-Xi Li, Ye-Hang Zhang, Hao-Jian Huang et al. · 0 citations
Preprint Aug 2026

SUN: Agentic Robot Policy Learning with Persistent Task Programs

Model-based control can directly execute specified objectives, while learning can amortize such behaviors into reactive policies, making their combination a natural solution to multi-stage manipulation. We introduce Semantically UNified (SUN) Programs, typed executables that compile grounded relations into aligned opti...

Wei-Qi Wang, Zhi Li, Yuliang Lei et al. · 0 citations
#natural language process... Preprint Sep 2026

Agent as Policy for Robotic Manipulation

This work demonstrates that a general-purpose agent can directly drive a physical robot throughout task execution without any task-specific or environment-specific training and introduces Agent as Policy (AGP), which places task planning and execution under the agent's control.

Meng-Zhao Jia, Yang Lin, Xi-Xin Zhang et al. · 7 citations · ⚡1
Preprint Sep 2026

RoboHarn-Evo: Evolving Hierarchical Physical Knowledge for Self-Improving Robotic Manipulation

Vision-language models can coordinate long-horizon robot manipulation, yet successful task reasoning still depends on whether local physical interactions produce the intended effects. We study how repeated interaction can improve this capability without updating the base model. We introduce RoboHarn-Evo, a dual-loop ha...

Shi-Feng Bao, Fan-Ding Huang, Yi-Han Lin et al. · 0 citations
Preprint Sep 2026

ASENA: Self-evolving Agents for Embodied Navigation

We present ASENA, an embodied agent system that connects general-purpose coding agents to robot sensing, computation, supervised execution, and persistent experience. Agents can write and execute programs, inspect recorded outcomes, repair failures, and reuse notes and executable skills while keeping their model weight...

An-Chieh Cheng, Isabella Liu, Edmund Bu et al. · 0 citations
Preprint Aug 2026

Compiling and Benchmarking Task-State Horizons for Embodied Agents

This work introduces RoboGraph, a robotic task compiler that translates state-transition dependencies into executable symbolic graphs, and constructs task-state horizons from spatial and temporal causal dependencies, including those induced by unexpected failures and interventions during task execution.

Mei Wang, Shi-Chao Li · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.