This paper introduces Physical Agentic AI, a framework for skill-grounded robot agent orchestration, in which each robot exposes a typed library of executable skills while a foundation model planner decomposes a task into phases and assigns each phase to a robot-skill pair.
Abstract
Agentic AI frameworks interpret open-ended task goals and decompose them into multi-step plans. Richer information about embodiment-specific capabilities, physical preconditions, and cross-robot coordination improves grounding, but does not eliminate infeasible, mistimed, or unsafe physical actions. Physical robot crews therefore require an explicit architectural interface between semantic planning and execution, where every planned action is verified against robot capabilities, system state, and workflow constraints before actuation. This paper introduces Physical Agentic AI, a framework for skill-grounded robot agent orchestration, in which each robot exposes a typed library of executable skills while a foundation model planner decomposes a task into phases and assigns each phase to a robot-skill pair. A Robot Orchestration layer exposes the skill library, robot state, named locations, and workflow contracts to a non-actuating Mission Planner, while a deterministic Robot Orchestrator validates and authorizes one skill at a time. We evaluate on a drone-UGV search-and-dispatch mission, where every mission in every condition is executed live in Gazebo, and on a humanoid-quadruped transportation task using hardware-equivalent skill interfaces plus two physical trials on a Unitree G1 and Go2. Varying planner knowledge and runtime enforcement independently, we find that retrieval raises skill grounding from 51% to 96% yet leaves informed planners dispatching 23-29% of faulted steps. Per-dispatch enforcement reduces false dispatch to 0% with no false blocks, and a held-plan ablation confirms that the gate, not plan variation, is responsible. Live execution makes the difference physical: without enforcement all eight injected faults crossed the orchestration boundary and six produced robot motion; with enforcement all eight were refused before motion.
These results validate the deployed navigation and inspection closed loop, while HROS provides an extensible software foundation for memory-augmented, voice-aware, and continuously improvable embodied inspection agents.
Yao-Yuan Yan, Zhi-You Heng, Hao-Xiang Jie et al.· 1 citation· ⚡1
Collective intelligence is a collaborative autonomy paradigm in which multiple agents pursue shared objectives through local perception, information exchange, and coordinated action. UAV swarms embody this paradigm by coordinating multiple vehicles in tasks such as search, inspection, and tracking. Recent advances in l...
Jiabin Lou, Yi-Rong Yang, Hao-Peng Wang et al.· 0 citations
This study presents a centralized safety aware M2M framework for cooperative goal-directed navigation in a heterogeneous mobile robot system composed of a vision-capable robot and a cameraless robotic vehicle that can approve safe motion, trigger conservative replanning or holding behavior, and preserve a strict separa...
Mohamed Dwedar, Ahmad Hafez, Alexander Jesser et al.· IEEE Access· 0 citations
AI planners can now generate a task plan for a team of robots, yet a human operator must stay in control of what the robots do. Doing so requires an interface that makes the plan visible and each step actionable, so mistakes can be caught before they reach the robots. Today, such an interface is developed for each spec...
Pattaraorn Yu, Ishaan Nair, Youngbin Song et al.· IEEE Robotics and Automation...· 0 citations
EmbodiedSkills is a unified framework that treats each skill decision as an execution proposal: the runtime checks its prerequisites before execution and verifies the outcome afterward and provides a trainable and inspectable agent layer for turning low-level VLA policies into closed-loop embodied systems.
Wei Wang, Wen-Qiao Zhang, Yu-Tong Lin et al.· 3 citations
This work formalizes the reasoning-execution boundary as a typed contract and constrains language-level decisions through schema-validated tool calls defined by the Model Context Protocol, rejecting malformed commands before they reach the robot.
Wen-Hao Hong, Lan Wei, Dan-Dan Zhang· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.