Skip to content
Open access

SRCDet: sparse fusion of surround-view radar and camera for 3D object detection

Aug 2026 · Measurement science and technology · Vol 37, pp. 365107 · 0 citations · 36 references
Physics

TL;DR

A query-based multimodal fusion framework, termed SRCDet, is proposed for camera-4D radar fusion, which achieves consistent improvements across nearly all metrics and low error rates in clear and adverse weather conditions, highlighting its practical adaptability to automotive-grade systems and effectiveness in safety-critical real-world autonomous driving scenarios.

Abstract

Multimodal fusion of cameras and millimeter-wave radars is critical for robust all-weather object detection and ensuring vehicle safety in real-world autonomous driving vehicles. However, existing radar-camera fusion methods that rely on a unified Bird’s Eye View (BEV) representation often suffer from information loss and limited cross-modal interaction. To address these limitations, a query-based multimodal fusion framework, termed SRCDet, is proposed for camera-4D radar fusion. The framework processes features in parallel across both BEV and Perspective View spaces, where deformable attention is employed to achieve dynamic cross-view alignment. By integrating radar attribute features, a local–global dual-branch query generation mechanism is designed to produce high-quality 3D detection proposals. Furthermore, a graph neural network-based cross-fusion module is introduced to model complex inter-feature relationships through a heterogeneous interaction graph. Extensive experiments on the OmniHD-Scenes and NuScenes datasets demonstrate that SRCDet achieves consistent improvements across nearly all metrics and low error rates in clear and adverse weather conditions, highlighting its practical adaptability to automotive-grade systems and effectiveness in safety-critical real-world autonomous driving scenarios.

Read PDF

Similar papers

2026

RC-HyFusion: Radar–Camera Hybrid Fusion for 3-D Object Detection in Autonomous Driving

In autonomous driving, achieving accurate and robust 3-D perception through the fusion of multiple sensor modalities is a critical requirement. While camera-based methods operating in the bird’s-eye view (BEV) have shown significant progress, they often suffer from performance degradation under adverse lighting and wea...

Li-Guo Chen, Yi-Peng Chen, Hong-Si Liu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

SARFusion: Scene-Aware Routing Fusion for Robust Camera-LiDAR 3D Object Detection

Camera-LiDAR fusion has become a prevailing paradigm for 3D object detection in autonomous driving. However, existing fusion detectors often establish strong inter-modality dependencies by decoding object queries from tightly coupled multimodal representations. Under corrupted driving conditions, such dependencies make...

Yu-Ting Zhao, Zi-Yi Zheng, Shu-Xiao Li · 0 citations
Nov 2026

DVFusion: Dual-View Attention and Geometry-Guided Fusion for 4D Radar-Camera 3D Detection

4D radar-camera fusion has attracted increasing attention for reliable 3D perception under adverse weather conditions. However, existing methods either lose valuable radar information by using sparse point clouds, or suffer from high computation when processing raw tensors directly. To address these challenges, this pa...

Li-Wei Luo, Xia Wu · 0 citations
Preprint Aug 2026

Boundary-Aligned Contribution Routing for Robust Optical--SAR Object Detection

Optical imagery provides rich appearance cues, whereas synthetic aperture radar (SAR) offers observations that are less sensitive to illumination and weather, making optical--SAR fusion attractive for remote-sensing object detection. However, the presence of multiple modalities does not guarantee beneficial fusion: imp...

Haifan Zhang, Yijing Wang, Haoyu Wang et al. · 0 citations
Aug 2026

Fadet: a fusion-aware 3D detection network with cascaded feature enhancement for small object detection in autonomous driving

This work proposes a cascade optimization framework that systematically enhances feature representation and refines multimodal fusion, and introduces the Multi-Scale Contextual Fusion Module (MSCF) to reduce alignment bias.

Chang-Hong Yu, Shaoshi Luo, Wen-Li Shen · 0 citations
Open access Aug 2026

MVXCC-NET: Cross-modal 3D detection of occluded objects based on dual-path information complementation and regional weight modeling

Surrounding scene awareness is a core component of self-driving techniques, and the 3D detection accuracy for occluded objects directly determines the system’s scenario adaptability and driving safety. To address the core problem of inter-object occlusions in traffic scenes that lead to reduced 3D detection accuracy an...

Jin Qi, Jian Wang · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.