Skip to content
Preprint

HODAgent: Towards On-Demand, Responsive Humanoids for Physical World Human Interaction

Aug 2026 · 0 citations · 60 references
Computer Science

TL;DR

Results show a unified System-2 agent enables adaptive humanoid service across simulation and reality.

Abstract

We propose HODAgent, a System-2 embodied agent for humanoid robots in service settings, addressing situated intent, responsive execution, task revision, and outcome verification. Its semi-duplex architecture integrates an Env-Interactor, Planner, Executor, and hierarchical Memory to maintain coherent interaction, planning, and task state during service episodes. This allows handling new requests during motion, retaining progress, revising actions, and grounding closure in execution outcomes. A shared interface connects simulation and physical robots (Unitree G1), isolating platform-specific control. In an interactive simulation with 164 cases, HODAgent achieves 84.8% and 91.5% Joint Success under two VLM backbones, outperforming baselines by 9.8 and 18.9 points. On physical robots, pass rates are 92% (atomic), 72% (composite), and 63.3% (complete tasks). On multiple embodied benchmarks, it improves over baselines by 0.7-9.0 points. Results show a unified System-2 agent enables adaptive humanoid service across simulation and reality.

View source

Similar papers

Preprint Jul 2026

WCM: World-Cognition Model for Generalizable Human-Robot Interaction

Language agents can now interact fluently with users in software, but robots still struggle to bring comparable interaction to physical tasks. Current robot-control paradigms, including vision-language-action policies and world-model-based planners, are mainly optimized for instruction execution, leaving users with little visibility into why an action is chosen and few mechanisms to redirect, correct, or teach the robot through interaction. To solve this problem, we present the World-Cognition Model (WCM), a human-centered embodied agent built on the SLAK architecture (Sensing, Logic, Action, and Knowledge) and an asynchronous runtime. SLAK separates perception, reasoning, control, and memory, while the runtime allows reasoning, dialogue, and execution to proceed concurrently. WCM further introduces a human-in-the-loop teaching mode that enables users to interactively teach the robot difficult or long-horizon tasks. Teaching episodes and autonomous task rollouts are refined into chain-of-thought supervision to continually improve the model. WCM achieves a 73.8% average success rate across nine real-world human-robot interaction tasks, including tasks held out from CoT fine-tuning and a long-horizon task learned through teaching.

Yuzhen Chen, K. Zhou · 0 citations
Preprint Aug 2026

ETA: A New Agentic Paradigm for Embodied Tasks

The Embodied Task Agent is introduced, a new paradigm for extending digital agents into the physical world, and OpenETA is released as its open-source implementation, which provides replaceable Planners, composable Tools and Skills, auditable memory, replayable trajectories, and common interfaces for simulation and real robots.

Yitong Chen, Zezheng Huai, Sixian Li et al. · 0 citations
Preprint Jul 2026

RoboBRIDGE: A Modular Framework for Bridging Policies to Robust Real-World Robotic Agents

Vision-Language-Action (VLA) models have attracted growing interest as a scalable approach to robotic manipulation. While these models are effective action predictors, deploying them as robotic agents exposes critical gaps: no mechanism for failure recovery, inconsistent execution over long horizons, and limited robustness to shifts in observations, tasks, or embodiments. Existing solutions address these limitations individually through model retraining or environment-specific modules, yet what is needed is a general framework that systematically transforms a pretrained VLA into a robotic agent. We present RoboBRIDGE, a modular framework that provides an orchestration layer over five coordinated modules, namely Monitor, Perceptor, Planner, Controller, and Robot Interface, to compose robust robotic agents from off-the-shelf components, including pretrained VLAs. The Monitor pairs rapid failure detection with hierarchical recovery to correct errors before they cascade. When the environment diverges from the current plan, the Planner triggers replanning while the Perceptor updates scene understanding asynchronously, avoiding execution stalls. Within the Controller, primitive skill fine-tuning factors manipulation into domain-invariant primitives with dedicated LoRA adapters, reducing sensitivity to domain shifts when a VLA is used. Across LIBERO, RoboCasa, and real-world case studies spanning multiple robot platforms and VLA backbones, RoboBRIDGE consistently outperforms both standalone policies and prior augmented VLA deployments. These results suggest that reliable robotic agency does not arise from scaling action predictors alone, but from structured orchestration around them.

Sihyung Yoon, Minjong Yoo, Sanghyun Ahn et al. · 1 citation
Open access Jul 2026

LLM-driven behavior generation and negotiation for collective robotic construction

Agent-based modeling and simulation (ABMS) has been widely employed to study emergent processes in collective robotic construction (CRC), where global architectural structures arise from local agent interactions. While these approaches reveal how complex assemblies can emerge without centralized control, they remain limited when an architectural goal is known but the behaviors required to achieve it are not. Most CRC workflows still depend on handcrafted heuristics. This paper presents a hybrid CRC workflow that integrates large language models (LLM) into the ABMS behavior design process. The system enables human–AI co-creation of robot behaviors, allowing an LLM agent to generate and negotiate behavioral strategies toward user-defined construction goals under partial observability. The approach is evaluated in simulation, comparing an LLM–heuristic hybrid against a heuristic-only baseline behavior. For well-documented swarm patterns, the LLM matches heuristic performance; for geometrically novel tasks, handcrafted heuristics retain an advantage. By embedding language-based reasoning within ABMS, this work expands participation in CRC behavior design and demonstrates a pathway for translating high-level design intent into adaptive, goal-oriented multiagent construction processes.

Samuel Slezák, Lasath Siriwardena, S. Leder et al. · 0 citations
Preprint Jul 2026

Courteous Anticipation: Improving Long-Lived Task Planning in Persistent Shared Environments

We consider a task planning scenario in which robots sharing a persistent environment are assigned tasks one at a time from a held-out sequence. Standard task planners, lacking foresight of future tasks and inconsiderate of others'constraints, solve each task in isolation, leaving terminal states that increase future cost for all, side effects that compound over lengthy task sequences. To reduce cost over the sequence, a robot must anticipate how its actions now may impact performance on future tasks for all robots sharing the environment. Therefore, we present courteous anticipatory planning, wherein a model-based planner proposes candidate plans and selects the one that jointly minimizes immediate cost and aggregated expected future cost across all robots, estimated via independent per-robot learned estimators. This factored formulation avoids combinatorial joint rollouts and supports modular deployment: adding a robot requires only training its own estimator. We evaluate in two persistent PDDL domains, a home environment with robots that have similar capabilities but different responsibilities, and a restaurant environment where robots'distinct capabilities create states that other robots lack the capability to resolve. During lengthy task sequences, our planner reduces total cost by 10.43% versus myopic and 4.03% versus selfish anticipatory planning in a two-robot home environment and by 17.41% and 13.24%, respectively, in a three-robot restaurant.

Md Ridwan Hossain Talukder, Roshan Dhakal, Elizabeth Phillips et al. · 0 citations