Skip to content
Preprint

Semantic Audio-driven Understanding for Dynamic Humanoid Whole Body Control

Jul 2026 · 0 citations · 33 references
Computer Science

TL;DR

This work introduces a novel multi-modal orchestration framework for semantic audio-driven humanoid control, enabling robots to autonomously select and execute appropriate motion skills in real time.

Abstract

Recent advances in humanoid robotics and reinforcement learning have enabled the acquisition of highly expressive whole-body motion policies. However, most robotic performances remain based on pre-scripted sequences or externally triggered behaviors, limiting autonomy and responsiveness to dynamic environments. In this work, we introduce a novel multi-modal orchestration framework for semantic audio-driven humanoid control, enabling robots to autonomously select and execute appropriate motion skills in real time. The system processes continuous audio streams and routes them into music or speech branches. Music input is handled via audio fingerprinting and semantic embeddings to retrieve track identity and temporal alignment, allowing dynamic mapping between musical segments and motion policies. Speech input is grounded into a discrete library of imitation-learned skills, enabling direct human-robot interaction. Both modalities share a unified interface that schedules skill execution over a reinforcement learning control pipeline. We validate the approach in simulation and on a Unitree G1 humanoid, showing robust sim-to-real transfer and consistent audio-conditioned policy selection. Supplementary materials are available at the following site: https://lab-rococo-sapienza.github.io/semantic-WBC/

View source

Similar papers

Preprint Aug 2026

Robot Learning from Human Demonstrations: Handwritten Alphabet Trajectories and Human-Likeness Evaluation

A framework for learning human-like robot motion from demonstration, including data collection, probabilistic trajectory learning, and perceptual user evaluation is presented, extending the widely used Gaussian Mixture Model and Gaussian Mixture Regression approach for learning from demonstration.

Alperen Kenan, Paul A. Bremner, Manuel Giuliani · 1 citation
Open access Jul 2026

Autoregressive Diffusion with Hybrid Representation for Interactive Human Motion Generation

ARDY is introduced, a streaming generation framework that bridges this gap by enabling high-fidelity motion generation controllable via online text prompts and flexible kinematic constraints and balancing precise trajectory control with efficient generative learning.

Kaifeng Zhao, Mathis Petrovich, Haotian Zhang et al. · 2 citations
Preprint Jul 2026

Learning Reusable Hybrid Motion Priors for Humanoid Locomotion from Motion Imitation

A three-stage pipeline that turns motion-imitation skills into a reusable hybrid motion prior (HMP) for humanoid locomotion and shows that training the codebook with the rotation trick improves latent organization and reduces downstream falls compared with a standard straight-through estimator.

Valerio Belli, Valerio Modugno, Enrico Mingo Hoffman et al. · 0 citations
Jul 2026

Continuous multi-skill motion generation for quadruped robots based on imitation–reinforcement learning

This work categorizes quadruped robot skills into three types: rhythmic motions, expressive motions, and high-dynamic motions and generates reference trajectories for each category using central pattern generators, animation design, and motion capture, respectively, and designs an asymmetric neural network architecture and employs an imitation–reinforcement learning algorithm to train policies for generating these three types of motions.

Chong Pi, Senwei Huang, Wei Li et al. · 0 citations