Skip to content
Conference

Speech Emotion Recognition Using Transformer-Based Architectures with Self-Attention Mechanisms

Aug 2026 · International Conference on Information Security and Cryptology · pp. 1295-1301 · 0 citations · 17 references

Abstract

Speech Emotion Recognition (SER) has become a key aspect in human-computer interaction, and affective computing, yet, current methods are faced with the challenge of modeling long-range context and fine-grain emotional expressions in speech signals. This paper has countered these shortcomings, giving a Transformer-based framework that incorporates a better self-attention to classify emotions better. The aim is to identify a solid and scalable SER architecture that can help detect various emotional states based on speech data. The procedure consists of combining multimodal acoustic feature extraction (MFCCs, Mel-spectrograms, and prosodic features) with Transformer-based modeling plus adaptive attention enhancement and contrastive learning techniques. Experiments on the Ryerson Audio-Visual Database of Emotional Speech and Song (RAVDESS) show better results, with a goal of 95.1% of accuracy and 94.8 of F1-score, surpassing a number of models at the state-of-the-art. These findings affirm that self-attention based architectures are highly effective at representing emotional features and classifying them. The paper concludes that the suggested strategy is a feasible and efficient method to utilize in any real-world SER application, and the improvements of the approach could be achieved by introducing multimodal and lightweight models.

View source

Similar papers

Open access Aug 2026

Intelligent Audio-based Emotion Recognition in Speech by Deep Learning and Feature Engineering Techniques

This work introduces ExpressNet, an optimum Multi-Layer Perceptron (MLP)-based SER model aimed to solve issues by leveraging a wide range of prosodic and spectral qualities incorporating Mel-Frequency Cepstral Coefficients (MFCCs), spectral contrast, and pitch variations.

Ramakrishna Gandi, A. Geetha, B. R. Reddy · 0 citations
Open access 2019

Speech Emotion Recognition using Convolutional Neural Networks and Recurrent Neural Networks with Attention Model

Speech emotion recognition is an upcoming subfield of automatic speech recognition that shares multiple similarities with mood recognition in music signals. Audio signals containing human speech are used as input to classification algorithms trained to recognize emotions in the form of audio features. This thesis outline...

G. Tomas, S. Weinzierl, Athanasios Lykartsis · 2 citations
Sep 2026

DeepAgeEmotionNet: A Multi-Task Deep Learning Framework for Speech-Based Age and Emotional State Assessment

The task of automatic speaker profiling based on speech signals becomes increasingly crucial in human-computer interaction, clinical voice assessment, and affective computing. Still, the tasks of speaker age group classification and emotion recognition are usually studied separately despite similar acoustic characteris...

R. Patole · 0 citations
Conference Open access Sep 2026

A Multimodal Speech Emotion Recognition Framework for Malayalam through Audio-Text Fusion

Speech emotion recognition (SER) is used in many domains, such as translation, intelligent assistants, healthcare monitoring, large language models, and human-computer interaction. Emotion recognition in Malayalam, however, remains challenging because of the language's rich morphological structure. This work introduces...

Athira Raj, Christy James Jose, K. Biju · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.