Adaptive Reliability-Guided Multimodal Learning Framework for Robust Emotion and Sentiment Classification Under Noisy and Missing Modalities
Abstract
Intelligent human–computer interaction, healthcare monitoring, social media analytics, and conversational AI rely on emotion recognition and sentiment analysis. However, multimodal learning systems frequently presume equal dependability between text, audio, and visual modalities, rendering them susceptible to environmental noise, missing information, and modality-specific errors. An Adaptive Reliability-Guided Multimodal Artificial Intelligence (ARGMAI) system for robust multimodal emotion and sentiment classification addresses these problems. In a single end-to-end architecture, reliability-aware preprocessing, adaptive cross-modal attention, dynamic feature fusion, contrastive representation learning, noise suppression, and cross-modal consistency optimisation are combined. Missing modality reconstruction allows the framework to preserve discriminative representations when one or more modalities are lacking. In model is tested using the public CMU-MOSEI and IEMOCAP datasets. Experimental results show that ARGMAI achieves 90.36% accuracy, 90.04% precision, 89.78% recall, 89.91% F1-score, and 96.08% AUC on the CMU-MOSEI dataset and 88.14% accuracy, 87.88% precision, 87.53% recall, 87.70% F1-score, and 88.41% weighted accuracy on the IEMOCAP The approach outperforms existing state-of-the-art multimodal learning algorithms under noisy and missing modality situations and provides consistent convergence and enhanced generalisation for real-world affective computing applications.