Skip to content

Understanding Multimodal Learning From Modality Fusion and Alignment Perspectives.

Sep 2026 · IEEE Transactions on Pattern Analysis and Machine Intelligence · Vol PP, pp. 1-15 · 0 citations
Medicine

TL;DR

This paper develops a dynamic strategy that jointly optimizes modality fusion and alignment, and develops a learning-based strategy using a bi-level optimization framework and theoretically proves the convergence of the learning algorithm to ensure its reliability.

Abstract

Despite significant advancements in multimodal learning (MML), it has been unexpectedly shown to underperform compared to unimodal approaches in practice, largely due to the modality imbalance problem, ultimately affecting the overall performance of the model. Naturally, most existing methods aim to rebalance optimization speeds across different modalities to avoid performance degeneration caused by modality imbalance. However, in addition to task-oriented modality fusion, we experimentally find that multimodal learning requires explicit modality alignment to stimulate weak modal capabilities so that they can be fully exploited, which is ignored by existing works. Therefore, in this paper, we explore the impact of modality fusion and alignment on multimodal learning from a unified perspective, and develops a dynamic strategy that jointly optimizes both, with particular emphasis on addressing modality imbalance. Concretely, we initially design a soft alignment strategy to impose the positive intervention from the prediction level by integrating modality fusion and alignment into a unified framework. We further extend this strategy to the representation level and hybrid level, enabling compatibility with a wider range of architectures. Subsequently, we design a heuristic strategy to dynamically integrate fusion and alignment. Furthermore, we develop a learning-based strategy using a bi-level optimization framework and theoretically prove the convergence of the learning algorithm to ensure its reliability. These two dynamic integration strategies are incorporated into a unified framework applicable to both supervised and semi-supervised scenarios, further enhancing performance. We conduct a series of experiments to demonstrate the effectiveness of our method on diverse datasets. Extensive experiments demonstrate that our method consistently outperforms state-of-the-art multimodal learning approaches, achieving accuracy improvements of 1.30%, 2.69%, and 0.60% on representative bimodal benchmarks, namely KSounds, CREMA-D, andSarcasm, respectively, as well as gains of 1.35% and 0.65% on trimodal datasets, namely NVGesture and IEMOCAP.

View source

Similar papers

Preprint Aug 2026

Mitigating Strong-Modality Collapse in Multimodal Learning via Inverted Asymmetric Fusion

Fusing multiple modalities is expected to improve model performance. However, on the MultiHuSE dataset, early, late, and symmetric attention fusion often fail to outperform the best unimodal baseline (text). Pathway isolation of a symmetric attention fusion model reveals that the text-pathway accuracy drops from 74.9%...

Mary Ogbuka Kenneth, Foaad Khosmood, Abbas Edalat · 0 citations
#machine learning Preprint Sep 2026

Confidence Falls Short: Asymmetric Certainty Gains from Optimization Hinder Multimodal Classification

Multimodal learning (MML) falls into the optimization dilemma due to the modality imbalance phenomenon, leading to suboptimal overall performance in practice. While many attempts primarily focus on balancing the optimization dynamics across modalities to address this issue, we identify a subtle yet critical flaw: optim...

Long-Fei Huang, Xiang-Yu Wu, Yang Yang · 0 citations
#artificial intelligence Preprint Sep 2026

Enabling Vision and Cross-Modal Learning for Multimodal Stroke Recurrence Prediction: An Interpretable Two-Step Framework

Multimodal stroke recurrence prediction requires effective integration of heterogeneous clinical and imaging data, yet modality imbalance often causes models to over-rely on dominant modalities and underutilize complementary information. While self-supervised pretraining and selective parameter freezing are commonly em...

Christian Gapp, Elias Tappeiner, Martin Welk et al. · 0 citations
Conference Aug 2026

Adaptive Cross-Modal Fusion With Instance-Level Gating for Vision-Language Understanding

Multimodal deep learning integrates heterogeneous data sources such as images and text to enable machines to understand complex real-world contexts. Although recent vision-language models have achieved significant progress, most existing approaches rely on rigid fusion strategies that combine modalities either at early...

Unnati A. Patel, Sanskruti Patel, J. Nanavati et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.