Skip to content

Author

J. García-Díaz

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Open access 2026

A Comparative Study of Audio-Language Models for Speech Emotion Recognition in Spanish

: Traditionally, speech emotion recognition has relied on supervised models that require task-specific training and annotated data. However, the recent emergence of audio-language models introduces a more flexible paradigm that enables multimodal reasoning through speech and natural language interaction. Nevertheless, their effectiveness for emotion recognition remains unclear. In this study, we evaluate audio-language models for speech emotion classification using the Spanish MEACorpus dataset and compare three approaches: prompt-based inference, embedding-based classification with lightweight classifiers, and instruction-tuned models with parameter-efficient fine-tuning plus a hybrid architecture based on class-specific confidence-driven routing. Our results show that the hybrid approach achieves the highest overall performance, reaching an 83.55% macro F1-score and an 84.37% weighted F1-score. Instruction tuning remains highly competitive, obtaining an 83.32% macro F1-score, which confirms the importance of supervised task adaptation for aligning ALMs with speech emotion recognition. We also include Whisper as a pretrained acoustic baseline to contextualize ALM-based representations against strong speech foundation models. Furthermore, the hybrid approach outperforms standalone embedding-based classification across all evaluated models, showing that class-specific confidence-driven routing can improve the use of embedding-based predictions. Although the best hybrid ALM configuration achieves competitive performance, it remains below the specialized MEACorpus baseline of 87.74% macro F1-score, indicating that general-purpose ALMs do not yet surpass highly specialized acoustic models for Spanish speech emotion recognition.

Jorge Gómez-Navalon, Ronghao Pan, Tomás Bernal-Beltrán et al. · 0 citations