Skip to content

Learning and Transferring Closed-Loop Robot Software

Sep 2026 · 0 citations · 25 references
Computer Science

TL;DR

The value of execution-improved software as a resource for acquiring new policies in this setting is demonstrated, although initial references remain better on two target tasks when averaged across runs.

Abstract

Closed-loop robot policies require observation processing, state management, and situation-dependent branching, making them costly to design and tune manually. Although coding agents increasingly support control-code generation and optimization, it remains unclear whether implementations improved on source tasks also support policy acquisition for new tasks. We study this question by treating complete closed-loop implementations as reusable execution experience. For each source task, a coding agent generates policy code from a few successful demonstrations and iteratively improves it using simulation feedback. The validation-selected implementations are retained in a software archive. For new tasks, the agent generates and improves policies using archived implementations, target demonstrations, and execution feedback. The resulting policy is then frozen and executes without further model calls. Across four source tasks in RoboCasa, iterative optimization increases mean success from 28.3% to 64.2%. Across nine target tasks and three independent runs, mean success is 45.2% without references, 41.5% with initial source code, and 57.0% with optimized source code. Optimized references outperform initial references in all three runs on the nine-task average, with a mean gain of 15.6 percentage points. These results demonstrate the value of execution-improved software as a resource for acquiring new policies in this setting, although initial references remain better on two target tasks when averaged across runs.

View source

Similar papers

#software testing Preprint Sep 2026

REVOLVE: An Automated Closed-Loop Framework for Evolving Robot Manipulation with Minimal Human Intervention

It is demonstrated that REVOLVE transforms real-world deployment into a closed-loop learning process that continually accumulates and uses execution experience, enabling continual evolution of both the policy and supervisory model with substantially less human intervention.

Han-Yu Liu, Qian Li, Yi-Zhu Ding et al. · 0 citations
Preprint Aug 2026

SUN: Agentic Robot Policy Learning with Persistent Task Programs

Model-based control can directly execute specified objectives, while learning can amortize such behaviors into reactive policies, making their combination a natural solution to multi-stage manipulation. We introduce Semantically UNified (SUN) Programs, typed executables that compile grounded relations into aligned opti...

Wei-Qi Wang, Zhi Li, Yuliang Lei et al. · 0 citations
#natural language process... Preprint Sep 2026

Agent as Policy for Robotic Manipulation

This work demonstrates that a general-purpose agent can directly drive a physical robot throughout task execution without any task-specific or environment-specific training and introduces Agent as Policy (AGP), which places task planning and execution under the agent's control.

Meng-Zhao Jia, Yang Lin, Xi-Xin Zhang et al. · 10 citations · ⚡1
#artificial intelligence Preprint Oct 2026

Reconstruct, Practice, Go Real: Guided Self-Improvement for Embodied Agents

Building reliable robot capabilities across diverse tasks requires substantial human effort to develop and maintain skills, design rewards, and integrate perception with control. We present Reconstruct, Practice, Go Real (RPG), a framework for autonomous improvement of robot execution systems without updating model wei...

Yen-Jen Wang, Hao-Zhe Jiang, Shu-Ying Deng et al. · 0 citations
#artificial intelligence Preprint Sep 2026

A Schema Bounded Language Model for Refining Robot Policies Without Destabilizing Local Learning

This paper addresses navigation by composite heterogeneous robots in a decentralized system when policy reasoning and local control operate at different update levels. In a NetLogo--Python implementation, three robots share motion dynamics but use different LLM backends. Each robot independently combines a large langua...

Chong-Wen Dong, Mithun Paul Saint-Germain, Pinjari Asif et al. · 0 citations
Preprint Sep 2026

EmbodiRSI: Recursive Self-Improvement for Data-Efficient Robot Adaptation

Adapting robot manipulation policies to new tasks and environments remains highly data-intensive, while the data needed for further improvement depends on the policy's current capabilities and failure modes. We introduce EmbodiRSI, an agentic system for recursive self-improvement (RSI) in a real-to-sim-to-real setting,...

Hao-Ran Lang, Hao-Tao Lu, Shi-Yu Sang et al. · 1 citation

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.