Skip to content
Conference Open access

A Multimodal Speech Emotion Recognition Framework for Malayalam through Audio-Text Fusion

Sep 2026 · Conference Proceedings in Science and Management · 0 citations · 3 references

Abstract

Speech emotion recognition (SER) is used in many domains, such as translation, intelligent assistants, healthcare monitoring, large language models, and human-computer interaction. Emotion recognition in Malayalam, however, remains challenging because of the language's rich morphological structure. This work introduces a multimodal SER model for Malayalam that uses both speech and text information. Because few Malayalam emotion databases are available, a dataset was constructed from publicly available audiovisual content. The method uses a wav2vec 2.0 model to extract audio features, while the text features are obtained with a pretrained language model. The unified model then fuses the audio and text features for emotion prediction. Label consistency in the constructed dataset was further examined with an embedding-based analysis. On a dataset of 1,040 samples evenly distributed across four emotions, the proposed model achieved an accuracy of 78.85% and a macro-averaged F1-score of 0.79 on the test set, outperforming an audio-only baseline.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.