Skip to content
Open access

Structure-Prior-Guided Multi-Stage Cross-Modal Collaborative Network for RGB-D Semantic Segmentation

Aug 2026 · Journal of Imaging · Vol 12, pp. 394 · 0 citations · 44 references
Medicine

TL;DR

The results suggest that separating reliability-oriented correction from multimodal fusion can limit the propagation of unreliable cross-modal responses and improve indoor RGB-D semantic segmentation performance.

Abstract

Red–green–blue and depth (RGB-D) semantic segmentation combines appearance cues from RGB images with geometric information from depth maps, but sensor noise, missing measurements, and boundary-inconsistent depth responses can introduce conflicting evidence during cross-modal fusion. We propose the Structure-Prior-Guided Network (SPGNet), a dual-branch, multi-stage framework that follows a correction-before-fusion strategy. At each feature scale, SPGNet estimates a learned structure prior from cross-modal agreement and discrepancy. The Cross-Modal Correction Module (CCM) uses this prior to regulate bidirectional information transfer, suppressing unreliable responses while retaining complementary cues. The Dual-branch Enhancement Fusion Module (DEF) then enhances the corrected RGB and depth features and integrates them through shared-representation-guided interaction, after which a lightweight multi-scale decoder produces the segmentation output. Under a unified training and evaluation protocol, SPGNet achieved three-run mean Intersection over Union (mIoU) scores of 50.845% on NYU Depth V2 and 48.457% on SUN RGB-D. Compared with the best reproduced baseline on each dataset, SPGNet improved mean mIoU by 2.111 and 0.899 percentage points, respectively. These results suggest that separating reliability-oriented correction from multimodal fusion can limit the propagation of unreliable cross-modal responses and improve indoor RGB-D semantic segmentation performance.

Read PDF

Similar papers

Open access Sep 2026

Boundary-Guided Dual-Perspective Cross-Modal Fusion Network for RGB-IR Object Detection

A Boundary-Guided Dual-Perspective Cross-Modal Fusion Network (BDPNet) is proposed to explicitly preserve shallow geometric structures and decouple deep semantic fusion into macroscopic and microscopic perspectives.

Hu Lin, Zhi-Wei Fu, Xiu-Mei Chen et al. · 0 citations
Open access Sep 2026

Orthogonal feature decoupling and hierarchical cross-scale aggregation network for RGB-T semantic segmentation

Visible-thermal (RGB-T) imaging systems provide crucial complementary information for robust visual sensing and measurement in complex illumination conditions. However, effectively fusing multi-modal sensory data to achieve accurate semantic segmentation remains a key challenge. Most existing methods rely on heuristic...

Hong-Wei Liu, Yi-Sha Liu, Wei-Min Xue et al. · 0 citations
Open access Sep 2026

BMF-DETR: Pseudo-Depth-Guided Bidirectional Multi-Strategy Fusion for End-to-End Object Detection

Transformer-based detectors model long-range context effectively, yet their representations remain dominated by RGB appearance and can become unreliable in cluttered, occluded, or crowded scenes. We present BMF-DETR, a pseudo-depth-guided detector that introduces RGB-derived geometric structure without requiring a dept...

Hai Wang, Jun-Hao Wen, Chun-Lai Yang et al. · 0 citations
Oct 2026

Robust Multimodal Gated Fusion for 3-D Object Detection via Alignment and Denoising

Multimodal perception integrating light detection and ranging (LiDAR) and cameras has become a key paradigm for 3-D object detection, as it leverages both geometric structure and semantic information. However, in real-world autonomous driving scenarios, calibration errors, adverse weather, and sensor degradation can in...

Hui-Lin Huang, Yan Bai, Peng-Yuan Wang et al. · 0 citations
Sep 2026

Edge-guided attention and multi-scale fusion for high-quality multi-modal image fusion

This work proposes EdgeAttnSwin, a framework that reformulates fusion as a three-stage progressive optimization process with explicit causal dependencies, and demonstrates that this method outperforms state-of-the-art approaches on infrared-visible and medical image fusion benchmarks.

Ke-Ni Tang, Ming-Yue Zhang, Hao-Ran Chen et al. · 0 citations
2026

D2CMFDet: Disparity-Guided Dynamic Cross-Modal Mamba Fusion for Multispectral Object Detection

Existing multispectral object detectors struggle to balance wide-area cross-modal discrepancy modeling with high-resolution local detail preservation. To address this bottleneck, we propose D2CMFDet, a disparity-guided dynamic fusion framework that integrates learnable state-transition modeling with cross-modal feature...

Hai Yang, Xing-Du Wu, Chen-Hai Wei et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.