Skip to content
Preprint

From Passive Execution to Active Exploration: Agentic Embodied Manipulation in Realistic Environments

Sep 2026 · 0 citations · 8 references
Computer Science

TL;DR

This work proposes an agent-based active exploration framework that enables robots to dynamically interact with the environment rather than merely execute predefined instructions, and introduces a fine-grained perception-execution interleaving strategy, which tightly couples visual feedback with skill execution to improve exploration robustness.

Abstract

Recent advances in agentic systems have substantially enhanced the long-horizon capability of embodied manipulation. However, many existing frameworks still follow a passive execution paradigm, which limits their applicability to real-world scenarios involving textual semantic cues, distractors, and initially invisible targets. To bridge this gap, we propose an agent-based active exploration framework that enables robots to dynamically interact with the environment rather than merely execute predefined instructions. Specifically, our framework consists of three collaborative modules: a planning module for high-level task reasoning, a perception module for visual scene understanding, and an execution module for low-level manipulation. This design allows the robot to actively acquire task-relevant information, adapt its behavior based on environmental feedback, and complete manipulation tasks under partial observability. Furthermore, we introduce a fine-grained perception-execution interleaving strategy, which tightly couples visual feedback with skill execution to improve exploration robustness. We evaluate our method on a realistic Find-and-Place task, demonstrating its effectiveness in challenging environments where target objects must be actively discovered before manipulation.

View source

Similar papers

Preprint Sep 2026

LIBERO-Agent: Evaluating General-Purpose Agents for Direct Embodied Manipulation

General-purpose agents can plan, use tools, and revise their behavior from feedback, but it remains unclear whether these capabilities transfer from digital environments to embodied manipulation. To investigate this question, we introduce LIBERO-Agent, an agent-native benchmark for evaluating these agents in robot mani...

Zi-Jie Diao, Yi-Tong Chen, Si-Cheng Xie et al. · 0 citations
Preprint Aug 2026

ETA: A New Agentic Paradigm for Embodied Tasks

The Embodied Task Agent is introduced, a new paradigm for extending digital agents into the physical world, and OpenETA is released as its open-source implementation, which provides replaceable Planners, composable Tools and Skills, auditable memory, replayable trajectories, and common interfaces for simulation and rea...

Yi-Tong Chen, Zezheng Huai, Si-Xian Li et al. · 5 citations · ⚡1
Preprint Sep 2026

RoboHarn-Evo: Evolving Hierarchical Physical Knowledge for Self-Improving Robotic Manipulation

Vision-language models can coordinate long-horizon robot manipulation, yet successful task reasoning still depends on whether local physical interactions produce the intended effects. We study how repeated interaction can improve this capability without updating the base model. We introduce RoboHarn-Evo, a dual-loop ha...

Shi-Feng Bao, Fan-Ding Huang, Yi-Han Lin et al. · 0 citations
#artificial intelligence Preprint Sep 2026

ActiveArena: Benchmarking and Understanding Active Perception in Robotic Manipulation

Active perception and manipulation are crucial for robots to interact with complex scenes. Existing benchmarks struggle to evaluate how robots effectively acquire and maintain information in memory in an active manner. To this end, we introduce ActiveArena-Sim, an active-perception simulator with controllable viewpoint...

Yi-Bo Li, En-Shen Zhou, Rui Chen et al. · 1 citation
Preprint Aug 2026

Zeva: In-Context Causal Learning for Generalizable Embodied Manipulation

Zeva is presented, the first framework that enables in-context learning from a robot's own physical interaction experience while keeping the policy model frozen, and achieves the best performance among the compared frontier VLAs and WAMs and enables self-evolution during deployment without gradient updates.

Fu Chen, Xin Ding, Bing-Jia Huang et al. · 3 citations
#artificial intelligence Preprint Sep 2026

Explore, Execute, Evolve: A Skill Acquisition and Reuse Loop for Embodied Agents

Vision-language-action and world-action models have demonstrated impressive capabilities in robotics, yet generalization to unseen tasks remains challenging. More recently, general-purpose multimodal agents have shown great potential for zero-shot robotic task solving. However, they often incur high execution costs by...

Si-Cheng Xie, Yi-Tong Chen, Hai-Dong Cao et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.