Skip to content

RoboFoundry: System-as-Policy Evolution for Self-Learning Embodied Agents

Sep 2026 · 0 citations · 26 references
Computer Science

TL;DR

RoboFoundry is proposed, the first embodied agentic framework that formulates this process of system-as-policy evolution across foundation models as Self-Evolving System-as-Policy, highlighting its potential for fully autonomous embodied agents.

Abstract

A foundation model should not act in isolation as an embodied agent. Yet, existing methods often optimize individual components of the agent stack, such as memory, context, skills, or action interfaces, rather than treating the supporting system itself as a unified policy. Moreover, interaction alone does not yield self-improvement unless execution experience is converted into persistent, validated system changes. We therefore propose RoboFoundry, the first embodied agentic framework that formulates this process as Self-Evolving System-as-Policy. RoboFoundry diagnoses capability gaps in decision-making and memory management, converts execution traces into validated task-specific system updates, and promotes recurring improvements to the general system. Evolution operates over two complementary surfaces: a context system that manages active internal context and persistent file-system memory, and a hierarchical skill system that organizes atomic skills, reusable compositions, and failure-conditioned recovery. A shared semantic interface separates embodiment-invariant decisions from embodiment-specific execution, allowing evolved system capabilities to transfer across heterogeneous robots. On EmbodiedBench, RoboFoundry achieves state-of-the-art performance, notably improving GPT-5.5 by 27.8%. It also brings Qwen3.7-Plus to near parity with GPT-5.5 (70.3% vs. 72.7%), showing consistent gains from system-as-policy evolution across foundation models. For long-horizon memory, RoboFoundry outperforms all baselines on RoboMemArena by at least 39.0%, even against methods assisted by external foundation models. On LIBERO-PRO, it further outperforms Cap-Agent0 by 243.8%-679.7% across all perturbation types. In real-world deployments, RoboFoundry demonstrates zero-shot transfer and online evolution across robots and tasks, highlighting its potential for fully autonomous embodied agents.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

MASkills: Continual Skills Optimization for Multi-Agent LLM Systems

MASkills presents a new agent-optimization pipeline that integrates skill-conditioned credit assignment, hierarchical credit aggregation, and momentum-smoothed optimization, enabling agent skill libraries to evolve through refinement, induction, consolidation, and pruning.

Huaiyuan Yao, Xiaoou Liu, Charles Fleming et al. · 1 citation
Preprint Sep 2026

RACaP: Agentic Reasoning, Acting, and Coding as Policies for Evolvable Robot Learning

General-purpose robot agents must learn from experience, transfer to new tasks, and act efficiently. Code as Policies (CaP) methods generate and repair programs at runtime, incurring latency and entangling reusable mechanisms with task-specific decisions. We introduce RACaP, an agentic framework that moves coding to ev...

Ze-Xi Li, Ye-Hang Zhang, Hao-Jian Huang et al. · 0 citations
#natural language process... Preprint Sep 2026

Experience Funnel: A State-Policy Alternating Loop for Self-Evolving Agents

Autonomous agents powered by large language models (LLMs) continuously accumulate experience through interaction, creating an opportunity to improve future behavior through self-evolution. A fundamental challenge is how to transform abundant, task-specific interaction experience into reusable model competence without s...

Wen-Bo Gao, Zhao-Mou Song, Zhi-Yuan Ji et al. · 0 citations
Preprint Aug 2026

EMERGE-Policy: A Robot Mind Emerges Beyond a Single Policy

A robot's effective ``mind''need not reside in a single policy. It can emerge when specialized components perceive, reason, predict, act, verify, and remember within a shared orchestration process. EMERGE-Policy turns this perspective into a graph-structured agentic framework that coordinates both capability invocation...

Zhile Fang, Qing-Chi Yu, Ziping Chen et al. · 0 citations
#artificial intelligence Preprint Sep 2026

DynaHarness: A Dynamic Physical Harness for Self-Evolving Robot Agents

Pretrained robot policies provide useful action priors, but long-horizon manipulation still requires coordination between semantic reasoning and physical execution. Semantic reasoning operates at a coarser timescale than physical interaction, while episode-level failures provide limited guidance on which system compone...

Hao-Yuan Deng, Jie-Bin Liu, Teng-Xiao Zhang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

AnyAct: Universal Action for Self-Evolving Agents

As large language models (LLMs) advance, AI agents are increasingly deployed in open-world environments to tackle complex sequential tasks (e.g., document processing, cross-application collaboration), relying heavily on actions ranging from GUI operations to semantic APIs. However, three core challenges persist: the"sc...

Ling-Rui Xu, Ya Jiang, Jia-Chang Zhang et al. · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.