This work introduces a novel multi-modal orchestration framework for semantic audio-driven humanoid control, enabling robots to autonomously select and execute appropriate motion skills in real time.
Abstract
Recent advances in humanoid robotics and reinforcement learning have enabled the acquisition of highly expressive whole-body motion policies. However, most robotic performances remain based on pre-scripted sequences or externally triggered behaviors, limiting autonomy and responsiveness to dynamic environments. In this work, we introduce a novel multi-modal orchestration framework for semantic audio-driven humanoid control, enabling robots to autonomously select and execute appropriate motion skills in real time. The system processes continuous audio streams and routes them into music or speech branches. Music input is handled via audio fingerprinting and semantic embeddings to retrieve track identity and temporal alignment, allowing dynamic mapping between musical segments and motion policies. Speech input is grounded into a discrete library of imitation-learned skills, enabling direct human-robot interaction. Both modalities share a unified interface that schedules skill execution over a reinforcement learning control pipeline. We validate the approach in simulation and on a Unitree G1 humanoid, showing robust sim-to-real transfer and consistent audio-conditioned policy selection. Supplementary materials are available at the following site: https://lab-rococo-sapienza.github.io/semantic-WBC/
A novel framework that leverages the reasoning capabilities of Large Language Models to synthesize complex robotic actions from a rich tapestry of multimodal human inputs: natural speech, hand gestures, and music/sound beats.
Snehasis Banerjee, R. Dasgupta· arXiv.org· 0 citations
A framework for learning human-like robot motion from demonstration, including data collection, probabilistic trajectory learning, and perceptual user evaluation is presented, extending the widely used Gaussian Mixture Model and Gaussian Mixture Regression approach for learning from demonstration.
Alperen Kenan, Paul A. Bremner, Manuel Giuliani· 1 citation
ARDY is introduced, a streaming generation framework that bridges this gap by enabling high-fidelity motion generation controllable via online text prompts and flexible kinematic constraints and balancing precise trajectory control with efficient generative learning.
Kaifeng Zhao, Mathis Petrovich, Haotian Zhang et al.· ACM Transactions on Graphics· 2 citations
A three-stage pipeline that turns motion-imitation skills into a reusable hybrid motion prior (HMP) for humanoid locomotion and shows that training the codebook with the rotation trick improves latent organization and reduces downstream falls compared with a standard straight-through estimator.
This work categorizes quadruped robot skills into three types: rhythmic motions, expressive motions, and high-dynamic motions and generates reference trajectories for each category using central pattern generators, animation design, and motion capture, respectively, and designs an asymmetric neural network architecture and employs an imitation–reinforcement learning algorithm to train policies for generating these three types of motions.
Chong Pi, Senwei Huang, Wei Li et al.· Robotica (Cambridge. Print)· 0 citations
The proposed RL-based dynamic control system successfully transforms score elements into performance actions which enable robots to deliver expressive music performances.