Skip to content

Agent as Policy for Robotic Manipulation

Sep 2026 · 7 citations · ⚡ 1 influential · 36 references
Computer Science

TL;DR

This work demonstrates that a general-purpose agent can directly drive a physical robot throughout task execution without any task-specific or environment-specific training and introduces Agent as Policy (AGP), which places task planning and execution under the agent's control.

Abstract

We demonstrate that a general-purpose agent can directly drive a physical robot throughout task execution without any task-specific or environment-specific training. We introduce Agent as Policy (AGP), which places task planning and execution under the agent's control. Given a task and a robot interface, the agent interprets visual evidence, writes executable programs, issues motion commands, and revises its actions in response to physical outcomes. This brings the agent's reasoning and programming capabilities into continuous interaction with the physical world. We study AGP across multiple real-world manipulation tasks spanning precision manipulation, dynamic motions, and deformable objects. These include assembly from human videos, block construction from goal images, dice flipping, targeted throwing, and bimanual towel folding. AGP achieves success rates of at least 80% in seven of eight task configurations and significantly outperforms previous agentic robot systems. We further study efficiency through task experience accumulation and find that reusing saved procedures and programs shortens execution time across repeated trials. These findings support a path for general-purpose agents to act as robot policies, extending their autonomy to physical manipulation through runtime reasoning, programming, and interaction.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

RAPID: Robot Agentic Programming from Demonstrations

This work introduces Robot Agentic Programming from Demonstrations (RAPID), which automatically generates, verifies, and refines robot programs, given a single visual human demonstration, using an object-centric relational program representation.

Yu-Yao Liu, Jia-Yuan Mao, David Hsu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

CodeActionBench: Evaluating Agentic Code-as-Policy for Embodied Manipulation

CodeActionBench is introduced, a benchmark of 25 manipulation tasks that evaluates this capability through agentic Code-as-Policy and provides a controlled testbed for measuring how general-purpose models translate their capabilities into manipulation behavior and for examining typical failure scenarios in that process...

Yiheng Lyu, Xueying Jiang, Wen-Hao Li et al. · 0 citations
Preprint Sep 2026

LIBERO-Agent: Evaluating General-Purpose Agents for Direct Embodied Manipulation

General-purpose agents can plan, use tools, and revise their behavior from feedback, but it remains unclear whether these capabilities transfer from digital environments to embodied manipulation. To investigate this question, we introduce LIBERO-Agent, an agent-native benchmark for evaluating these agents in robot mani...

Zi-Jie Diao, Yi-Tong Chen, Si-Cheng Xie et al. · 0 citations
Preprint Aug 2026

Revisiting the"Push-T"Robot Manipulation Task with Agentic Robotics

This short paper revisits the Push-T task in the context of emerging advances in Agentic Robotics where an LLM coding agent -- Claude Code with Fable 5 -- is prompted to create an algorithmic solution that does not require any demonstration data.

Shuang-Yu Xie, Kai-Peng Chen, Ken Goldberg · 0 citations
Preprint Aug 2026

SUN: Agentic Robot Policy Learning with Persistent Task Programs

Model-based control can directly execute specified objectives, while learning can amortize such behaviors into reactive policies, making their combination a natural solution to multi-stage manipulation. We introduce Semantically UNified (SUN) Programs, typed executables that compile grounded relations into aligned opti...

Wei-Qi Wang, Zhi Li, Yuliang Lei et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.