Skip to content

RiVaT-Fuse: Reliability-Calibrated Variational Tensor Fusion for Multimodal Prediction under Modality Uncertainty

Sep 2026 · 1 citation · 45 references
Computer Science

TL;DR

RiVaT-Fuse is proposed, a reliability-calibrated variational tensor fusion framework that defines fusion as sample-wise latent-state estimation and achieves the strongest overall predictive rank among direct representation-level baselines while improving probability and label stability under perturbation.

Abstract

Image-metadata prediction requires fusing heterogeneous evidence whose reliability can vary across samples and latent factors. Existing representation-level fusion methods typically choose an aggregation architecture, such as concatenation, gating, conditional modulation, or attention, without explicitly defining what the fused representation should mean under modality uncertainty. We propose RiVaT-Fuse, a reliability-calibrated variational tensor fusion framework that defines fusion as sample-wise latent-state estimation. Rather than producing a fused vector by direct aggregation, RiVaT-Fuse estimates a consensus latent state through a variational objective that balances image evidence, metadata evidence, structured cross-modal interaction, and stability. The resulting framework replaces scalar modality confidence with matrix-valued trust geometry, decomposes interaction into additive, multiplicative, and relational components, and couples the latent state with conditional robustness and structured multi-task prediction. We provide well-posedness and stability interpretations of the latent solve and instantiate the framework with efficient low-rank-plus-diagonal trust operators. On an image-level image-metadata prediction benchmark, RiVaT-Fuse achieves the strongest overall predictive rank among direct representation-level baselines while improving probability and label stability under perturbation.

View source

Similar papers

#machine learning Preprint Sep 2026

Structured Latent Modeling for Supervised Multimodal Information Decomposition

Multimodal prediction relies on diverse forms of evidence: information repeated across modalities, cues specific to a single source, and complex cross-modal dependencies that emerge only when inputs are considered together. While recent methods promote richer interactions, they lack a principled way to isolate these ta...

Wan-Ting Huang, Sanvesh Srivastava, Wei-Ran Wang · 0 citations
Preprint Aug 2026

Conformal Fusion Under Missing Modalities

Modality-Conditioned Conformal Fusion is, to the authors' knowledge, the first method with formal coverage guarantees under arbitrary modality availability through architectural integration rather than post-hoc recalibration, and the evidential decomposition yields per-modality vacuity scores that localise uncertainty...

Alireza Moayedikia · 0 citations
Preprint Sep 2026

Beyond Spatial-Domain Supervision: A Relation Constrained Space for Multi-Modal Image Fusion

Multi-modal image fusion (MMIF) aims to form a single image by integrating shared information, preserving complementary cues, and coordinating cross-modal conflicts across modalities. However, due to the absence of ground-truth fused images, existing MMIF supervision commonly uses spatial-domain sources or gradient var...

Ze-Yu Wang, Ming-Yu Ge, Hai-Yu Song et al. · 0 citations
Preprint Sep 2026

Refine Then Fusion: Training-Free 3D Point Cloud Adaptation with Priority Refinement and Multi-Modal Knowledge Fusion

Recent pre-trained foundation models provide rich multi-modal priors for downstream 3D vision tasks. However, the effectiveness of these representations in few-shot scenarios is limited by two fundamental challenges: High-dimensional features often contain substantial channel redundancy and task-irrelevant noise, while...

Hang Cheng, Yan Chen, Ming-Yu Fan et al. · 0 citations
Open access Sep 2026

A unified multi-modal latent diffusion framework with modality-dropout training

Introduction Multi-modal conditioning in latent diffusion models—combining text, structural, and spatial guidance signals—substantially improves controllable image synthesis, yet two limitations persist across most existing frameworks. First, conditioning modalities are fused using fixed architectural weights that rema...

S. Remya, Manu J. Pillai, Laveena Herman et al. · 0 citations
Conference Open access Sep 2026

Trust, but Verify: Uncertainty-Driven Evidential Multimodal Representation Learning

This work proposes Adaptive Evidential Multimodal Representation Learning (AEMRL), a framework aligned with this multi-level view of uncertainty that consistently enhances task accuracy, reduces calibration error, and improves robustness to noise, semantic conflict, and missing modalities.

Yu-Peng Han, Kai Zhang, Xian-Quan Wang et al. · 0 citations

Related blog posts

Microsoft Research Blog Aug 11, 2026

Introducing CARE-X: Towards Clinically Useful Radiology VLMs with Auxiliary Supervision, Reward-Aligned Learning, and Tool-Augmented Measurement

Radiology AI is evolving beyond report generation. CARE-X explores a unified approach that combines flexible reasoning, calibrated predictions, and measurement-based tools for chest X-ray interpretation. The post Introducing CARE-X: Towards Clinically Useful Radiology VLMs with Auxiliary Supervision, Reward-Aligned Learning, and Tool-Augmented Measurement appeared first on Microsoft Research.

MIT News · Artificial Intelligence Oct 7, 2026

Discovering the value of humanistic inquiry

Students in MIT’s Concourse program delve deeply into the human condition, debate challenging questions, and learn to develop judgment about issues that can’t be quantified.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.