Skip to content

Author

Kalyan Thapa

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Conference Jul 2026

Multimodal AI Framework for Music Emotion and Media Sentiment Analysis using Deep Learning

With the evolution of multimedia platforms and online streaming services, the need for intelligent systems have increased to learn from heterogeneous data sources to understand human emotions and sentiments. Conventional unimodal paradigm for music emotion recognition and sentiment analysis, which relies solely on the use of single mode data (usually audio modality), has showed limited success in the task. Towards overcoming these challenges, this paper introduces an innovative cross-modal transformer based multimodal deep learning framework for combined music emotion and media sentiment analysis using synchronized multimodal representations from the CMU-MOSEI dataset. It proposes a framework which deploys Convolutional Neural Network (CNN) and Bidirectional Long Short-Term Memory (BiLSTM) models to extract music-inspired acoustic emotional features, Bidirectional Encoder Representations from Transformers (BERT) for textual sentiment representation learning, and Vision Transformer (ViT)-based feature extraction for visual emotional understanding respectively. The cross-modal transformer attention fuses heterogeneous modality representations and enhances the contextual interaction learning. Experimental evaluation shows that the proposed framework significantly outperforms traditional unimodal and multimodal approaches in terms of accuracy, precision, recall and F1-score. The proposed system is a solid and scalable solution for next generation affective multimedia analytics, intelligent recommendation systems and emotion-aware digital media applications.

Madhur Thapliyal, Anuja Rohilla, Ashish Kulshrestha et al. · 0 citations