Skip to content
Conference

Adaptive Cross-Modal Fusion With Instance-Level Gating for Vision-Language Understanding

Aug 2026 · International Conference on Circuit, Power and Computing Technologies · pp. 425-432 · 0 citations · 20 references

Abstract

Multimodal deep learning integrates heterogeneous data sources such as images and text to enable machines to understand complex real-world contexts. Although recent vision-language models have achieved significant progress, most existing approaches rely on rigid fusion strategies that combine modalities either at early or late stages without dynamically adjusting modality contributions. Such strategies may become unreliable when one modality is noisy, incomplete, or less informative.This paper proposes an adaptive attention-based fusion framework that dynamically balances visual and textual representations. The proposed architecture incorporates bidirectional cross-modal attention together with an instance-level gating mechanism that determines the relative importance of each modality for every input sample. Visual and textual features are first extracted using modality-specific encoders and then aligned through cross-modal attention layers. Subsequently, a lightweight gating network assigns modality weights to construct an interpretable fused representationThe training objective integrates task-specific supervision, contrastive alignment, and regularization to prevent modality dominance. Experiments on VQA v2 and MS-COCO benchmarks demonstrate consistent improvements over static fusion approaches and representative vision-language transformer baselines. The proposed framework also exhibits robustness when one modality is missing and provides interpretable modality contributions.

View source