Skip to content
Conference

Emotion-driven music retrieval via cross-modal audio-visual representation learning

Aug 2026 · International Conference on Artificial Intelligence, Big Data and Electrical Automation · Vol 14319, pp. 143191E - 143191E-10 · 0 citations · 17 references
Engineering

Abstract

Music-assisted therapy relies on the accurate perception of a user's affective state to deliver contextually appropriate acoustic stimuli. However, existing emotion-driven retrieval methods frequently encounter representational mismatches between raw multimodal behavioral signals and therapeutic music indexing, which stem from the limitations of unimodal sensing and linear feature fusion strategies in complex environments. To mitigate these challenges, a cross-modal audio-visual representation learning framework (CAV-RL) is proposed, which integrates multimodal feature extraction, cross-modal interaction, and affective mapping to enable emotion-aware music retrieval. Specifically, the sensing module employs a dual-branch architecture, integrating a self-supervised Wav2Vec 2.0 acoustic encoder with a spatiotemporal visual Transformer to extract prosodic variations and facial micro-expressions. A Cross-Modal Attention (CMA) mechanism is implemented to explicitly model asynchronous cross-modal dependencies, enabling the model to attend to visually salient regions guided by auditory cues. In addition, an adaptive confidence gating strategy modulates modality-specific contributions, enhancing discriminative robustness against isolated signal degradation. Moreover, to mitigate the semantic gap between discrete categorization and continuous retrieval, a probabilistic expectation mapping mechanism is introduced. The predicted categorical distributions are mathematically projected into a continuous Valence-Arousal (V-A) coordinate space, facilitating zero-shot, distance-based music retrieval without the necessity for cross-domain retraining. Empirical evaluations on the RAVDESS dataset demonstrate that the proposed framework achieves competitive classification accuracy, which further enables the effective translation of multimodal emotional states into quantifiable retrieval metrics.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.