2026· ITM Web of Conferences· 0 citations· 10 references
TL;DR
This review offers an all-round synthesis of how LiDAR and camera sensors are integrated for 3D target detection and aims to give a comprehensive theoretical explanation to scholars who have a preliminary understanding of the fusion of LiDAR and camera.
Abstract
This review offers an all-round synthesis of how LiDAR and camera sensors are integrated for 3D target detection. It first outlines why LiDAR–camera fusion is necessary and where its advantages lie, then turns to the respective characteristics and underlying theoretical bases of LiDAR and cameras, particularly with respect to detection range, measurement accuracy, and semantic richness. From there, the discussion examines the operating principles of both sensing modalities and, more to the point, the central challenge in fusing them: the heterogeneity of the data each system acquires. Against that technical backdrop, the analysis proceeds to the current mainstream approaches to feature-level fusion. Through feature- level fusion, LiDAR point cloud features and camera image features can be effectively integrated, which broadens the environmental perception scope and supports accurate extrinsic calibration for reliable measurement. Finally, this paper further analyzes the existing problems in this field and prospects the future development direction. This paper aims to give a comprehensive theoretical explanation to scholars who have a preliminary understanding of the fusion of LiDAR and camera.
A reliability-gated cross-modal fusion model, RG- CMF, is proposed for complex-scene 3D object recognition that integrates calibration-aware RGB–LiDAR alignment, dual- branch feature extraction, reliability-gated fusion, and geometry–semantic consistency optimization and suggests that explicit reliability modeling can...
Zhongxing Duan· International Conference on...· 0 citations
In autonomous driving, achieving accurate and robust 3-D perception through the fusion of multiple sensor modalities is a critical requirement. While camera-based methods operating in the bird’s-eye view (BEV) have shown significant progress, they often suffer from performance degradation under adverse lighting and wea...
Li-Guo Chen, Yi-Peng Chen, Hong-Si Liu et al.· IEEE Transactions on Aerospa...· 0 citations
To address the challenges of insufficient feature representation and the difficulty of detecting sparse and distant objects in UAV-borne LiDAR point clouds—which exhibit significantly lower point density than terrestrial/mobile LiDAR scans—this paper proposes an enhanced detection algorithm built upon the PointPillars...
Yu Zhai, Sen Xie, Wen-Hao Li et al.· Electronics· 0 citations
A target-free wide-area calibration method for roadside LiDAR-camera systems that estimates extrinsic parameters directly from natural traffic scenes, providing higher calibration accuracy while preserving practical computational efficiency and showing good robustness under challenging conditions.
The rapid advancement of autonomous driving and embodied intelligence highlights the critical need for robust multi-sensor systems. In such systems, LiDAR and cameras are frequently paired—leveraging their complementary strengths: LiDAR excels at high-precision measurements, while cameras deliver rich texture informati...
Bao-Sheng Zhang, Lin Zhang, Sheng-Jie Zhao et al.· IEEE Transactions on Image P...· 0 citations
This work proposes a cascade optimization framework that systematically enhances feature representation and refines multimodal fusion, and introduces the Multi-Scale Contextual Fusion Module (MSCF) to reduce alignment bias.