The proposed RL-based dynamic control system successfully transforms score elements into performance actions which enable robots to deliver expressive music performances.
Abstract
Music performance enables robots to show their emotions to humans through musical expression. Traditional systems face difficulties when they need to adapt to changing musical expressions and performance environments. The purpose of this research is to develop a dynamic control strategy for a violin-playing robotic system using Reinforcement learning (RL) to improve expressive performance. The proposed approach adapts key bowing dynamics, including bowing speed, bow pressure, bow direction, and timing, based on musical cues. An Entropy-regularized radial basis function with deep Q-network (ER-RBF-DQNet) model is introduced as the core framework. This hybrid RL architecture enhances nonlinear feature mapping and adaptive decision-making. It also improves exploration for more expressive robotic violin performance. A dataset of symbolic musical scores labeled with pitch, duration, and target sound pressure served as input. Normalization techniques were applied to scale musical features, and noise filtering was used to remove inconsistencies in dynamic annotations. Mel-frequency cepstral coefficients (MFCCs), along with tempo, pitch contour, and dynamic range, were extracted as expressive control parameters. The proposed method combines RBFN, DQN, and ER to map musical input features, approximate optimal Q-values, and encourage exploration during training. This allows the robot to generate control signals that align with the musical score, enhancing human-robot interaction through music. Python was implemented, and the evaluation, the accuracy (98.6%) indicates the proportion of robotic control operations that accurately matched the annotated ground-truth musical score. The proposed RL-based dynamic control system successfully transforms score elements into performance actions which enable robots to deliver expressive music performances.
This work introduces a novel multi-modal orchestration framework for semantic audio-driven humanoid control, enabling robots to autonomously select and execute appropriate motion skills in real time.
J. Marcelo, M. Brienza, E. Bugli et al.· 0 citations
Introduction The cognitive encoding of musical sequences is a complex process that involves capturing the intricate structure, temporal dynamics, and inherent uncertainties of musical data. Traditional methods often struggle to preserve the non-Euclidean geometric properties of musical sequences and fail to adequately model temporal dependencies and uncertainties. This paper introduces the Manifold Adaptive Sequence Encoder (MASE), a novel neural framework designed to address these challenges. Methods MASE integrates three key modules: the Riemannian Trajectory Mapper, which embeds musical sequences into a Riemannian manifold to maintain their geometric properties; the Agent-driven Temporal Planner, which effectively models the temporal dependencies and rhythmic patterns; and the Uncertainty-guided Sequence Filter, which quantifies and incorporates uncertainty to enhance robustness and generalization. The framework is optimized using manifold alignment optimization, ensuring the alignment of latent representations with the input data, and uncertainty-aware refinement, which iteratively refines predictions by leveraging uncertainty estimates. Results and discussion Experimental results demonstrate that MASE significantly improves the accuracy and robustness of musical sequence modeling, outperforming existing methods by a substantial margin. The proposed approach offers a principled methodology for modeling the cognitive encoding of musical sequences, with potential applications in music analysis, recommendation, and generation. This advancement in musical sequence encoding not only enhances the understanding of cognitive processes involved in music perception but also provides a robust tool for various practical applications in the field of music technology.
Shao-Jie Lin, Guang Zeng· Frontiers in Psychology· 0 citations
Deep learning has been widely applied to digital art and music creation. However, producing melodies that follow music theory and match human compositional patterns remains challenging. This study proposes a symbolic music generation system that integrates supervised learning and reinforcement learning. The core framework employs recurrent neural networks for sequence modeling, while a reinforcement learning module formulates music rules as reward functions to guide pitch, duration, and rhythm. This hybrid approach helps produce outputs that better match human compositional patterns. We deploy the proposed framework as an Internet-based system that supports five distinct music styles and is publicly accessible online.
C. Wu, Timothy K. Shih, Chien-Hao Huang et al.· Journal of Internet Technolo...· 0 citations
Advances in technology have led to increasingly sophisticated musical humanoid robots. However, their use has largely been limited to performance and related research in human-robot interaction. In this position paper, we propose a novel perspective: musical humanoid robots as experimental interfaces for investigating music-evoked emotions. We argue that current research is constrained by paradigms relying on pre-recorded auditory stimuli, which fail to capture the multimodal, embodied, and interactive nature of real-world musical experience. Building on existing theories of music cognition and emotion, we identify mechanisms that require controlled manipulation of both acoustic and non-acoustic variables. We show that humanoid robots are well-suited as they enable parametric control of performance variables, reproducibility across trials, and the decoupling and recombination of auditory, visual, and interactive components. We illustrate the technical feasibility of this perspective through a case study of the WAseda Saxophonist Robot 5 (WAS-5), demonstrating reproducible control of acoustic and interaction variables that are prerequisites for future music-emotion experiments. Our work positions musical humanoid robots as a methodological platform that enables future controlled investigations of music-evoked emotions.