Skip to content

Visual Mapping of Multimodal Emotion Features for Intelligent Interaction Design.

Aug 2026 · Journal of Visualized Experiments · Vol 234 · 0 citations
Medicine

Abstract

By enabling machines to comprehend, interpret, and react to human affective states, visual mapping of multimodal emotion traits is crucial to intelligent interface design. Nevertheless, current multimodal emotion detection algorithms primarily focus on predictive performance, providing little insight into interpretability, interaction awareness, or feature-level cross-modal interactions. This work introduced a framework for the visual mapping of multimodal emotion aspects that combines interaction-oriented visualization, interpretable representation learning, and emotion prediction. Without using manually created features, transformer-based encoders were used to extract high-level emotion embeddings from textual, audio, and visual modalities. To improve robustness, emotion-specific and modality-specific information were separated using a disentangled representation learning technique. Additionally, a graph-based cross-modal attention network was created to use adaptive attention weighting to simulate subtle emotional relationships between modalities. A multi-view visual mapping module was created to depict temporal emotion trajectories, cross-modal attention distributions, and latent emotion embeddings in order to improve transparency. The Carnegie Mellon University Multimodal Opinion Sentiment and Emotion Intensity (CMU-MOSEI) dataset and the Chinese Multimodal Sentiment Analysis Dataset (CH-SIMS) were used to assess the suggested framework. According to experimental results, the suggested framework provided improved feature-level interpretability and interaction-aware visualization while consistently outperforming representative baseline techniques in terms of accuracy, F1-score, correlation, and mean absolute error. In addition, the visual mapping module facilitated qualitative interpretation of multimodal emotion representations, cross-modal interactions, and temporal emotion dynamics, supporting interaction-aware visualization and analysis.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.