CAMMRF: A Conflict-Aware Missing-Modality Robust Fusion Framework for Multimodal Sentiment Analysis under Heterogeneous Observability
Abstract
Multimodal sentiment analysis combines linguistic, acoustic, and visual evidence, yet deployment commonly violates the benchmark assumption that every modality is present and mutually consistent. Missing-modality systems generally reconstruct absent channels while conflict-aware systems generally assume full observation, even though absence and disagreement both change how far an observed modality should be trusted. We therefore treat the two jointly and hypothesise that conditioning modality reliability on both availability and cross-modal compatibility, and applying it across aligned and disagreement-preserving representations, yields more consistent performance across observability regimes than matched fusion controls without this coupling. We introduce CAMMRF, a framework that couples a mask- and peer-conditioned reliability estimator, an alignment/conflict subspace decomposition, and stochastic modality dropout with cross-view consistency. We evaluate one frozen protocol on the official CMU-MOSI train/validation/test partitions under eight observability regimes; five independent training seeds (42, 43, 44, 45, and 46) completed, each retaining predictions, checkpoints, and logs. In the full-modality regime CAMMRF is best on all four metrics simultaneously—MAE 1.050 ± 0.028, Pearson correlation 0.593 ± 0.009, weighted F1 75.01 ± 1.72%, and Acc-2 74.91 ± 1.80%—rather than on accuracy alone, and it ranks first by mean Acc-2 in 5 of eight regimes. The advantage persists when audio or vision is missing and under the controlled conflict stress (mean Acc-2 73.81%, a 1.10-point decrease, with matching small changes in MAE, weighted F1, and correlation); however, all four metrics degrade together when text is absent—correlation most steeply—so robustness is regime-dependent rather than universal. Full-budget ablations show that several component removals improve full-observation point estimates, supporting an accuracy–robustness trade-off rather than an independent-component claim. The evidence is limited to one corpus and controlled surrogate baselines and does not establish external generalisation. A single Colab notebook regenerates the protocol and exports a machine-readable result registry.