Aug 2026· Measurement science and technology· Vol 37, pp. 365107· 0 citations· 36 references
Physics
TL;DR
A query-based multimodal fusion framework, termed SRCDet, is proposed for camera-4D radar fusion, which achieves consistent improvements across nearly all metrics and low error rates in clear and adverse weather conditions, highlighting its practical adaptability to automotive-grade systems and effectiveness in safety-critical real-world autonomous driving scenarios.
Abstract
Multimodal fusion of cameras and millimeter-wave radars is critical for robust all-weather object detection and ensuring vehicle safety in real-world autonomous driving vehicles. However, existing radar-camera fusion methods that rely on a unified Bird’s Eye View (BEV) representation often suffer from information loss and limited cross-modal interaction. To address these limitations, a query-based multimodal fusion framework, termed SRCDet, is proposed for camera-4D radar fusion. The framework processes features in parallel across both BEV and Perspective View spaces, where deformable attention is employed to achieve dynamic cross-view alignment. By integrating radar attribute features, a local–global dual-branch query generation mechanism is designed to produce high-quality 3D detection proposals. Furthermore, a graph neural network-based cross-fusion module is introduced to model complex inter-feature relationships through a heterogeneous interaction graph. Extensive experiments on the OmniHD-Scenes and NuScenes datasets demonstrate that SRCDet achieves consistent improvements across nearly all metrics and low error rates in clear and adverse weather conditions, highlighting its practical adaptability to automotive-grade systems and effectiveness in safety-critical real-world autonomous driving scenarios.
In autonomous driving, achieving accurate and robust 3-D perception through the fusion of multiple sensor modalities is a critical requirement. While camera-based methods operating in the bird’s-eye view (BEV) have shown significant progress, they often suffer from performance degradation under adverse lighting and wea...
Li-Guo Chen, Yi-Peng Chen, Hong-Si Liu et al.· IEEE Transactions on Aerospa...· 0 citations
Camera-LiDAR fusion has become a prevailing paradigm for 3D object detection in autonomous driving. However, existing fusion detectors often establish strong inter-modality dependencies by decoding object queries from tightly coupled multimodal representations. Under corrupted driving conditions, such dependencies make...
4D radar-camera fusion has attracted increasing attention for reliable 3D perception under adverse weather conditions. However, existing methods either lose valuable radar information by using sparse point clouds, or suffer from high computation when processing raw tensors directly. To address these challenges, this pa...
Li-Wei Luo, Xia Wu· IEEE Robotics and Automation...· 0 citations
Optical imagery provides rich appearance cues, whereas synthetic aperture radar (SAR) offers observations that are less sensitive to illumination and weather, making optical--SAR fusion attractive for remote-sensing object detection. However, the presence of multiple modalities does not guarantee beneficial fusion: imp...
Haifan Zhang, Yijing Wang, Haoyu Wang et al.· 0 citations
This work proposes a cascade optimization framework that systematically enhances feature representation and refines multimodal fusion, and introduces the Multi-Scale Contextual Fusion Module (MSCF) to reduce alignment bias.
Surrounding scene awareness is a core component of self-driving techniques, and the 3D detection accuracy for occluded objects directly determines the system’s scenario adaptability and driving safety. To address the core problem of inter-object occlusions in traffic scenes that lead to reduced 3D detection accuracy an...
Jin Qi, Jian Wang· Journal of King Saud Univers...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.