Practical Human–Robot Interaction (HRI) through Large Language Model (LLM)-based Voice-to-Action Systems
Voice interaction has been actively studied in human–robot interaction (HRI) for decades, yet deploying spoken interfaces on physical mobile manipulators remains challenging because language is ambiguous, tasks are long-horizon, and robot actions must be grounded to perception and motion under real-world uncertainties. Recent large language models (LLMs) offer a practical way to interpret open-ended spoken requests, but their non-deterministic outputs and limited transparency can hinder safe and reproducible robot execution. This paper presents an LLM-based Voice-to-Action (VTA) system that converts spoken user commands into grounded robot behaviors for an indoor mobile manipulator, LeeAhn 2. The system combines speech transcription with an LLM that produces structured, skill-level action plans aligned with a predefined library of robot capabilities, including vision-based seeking, wheeled navigation, and manipulation. To improve reliability, we incorporate interface constraints that restrict generated actions to executable skills and enable recovery from common failures during execution. We evaluate the proposed system through simulation experiments and real-world trials on a representative search task, reporting component-level performance across seeking, navigation, and manipulation. The simulation experiments provide repeatable analysis under controlled conditions, while the real-world trials demonstrate practical applicability on the LeeAhn 2 mobile manipulator and reveal limitations such as latency and occasional plan/execution failures. Our results suggest that LLM-grounded spoken interfaces can reduce operator burden and improve accessibility for indoor service robots.