Skip to content
Conference

Intelligent 3D object recognition in complex scenes based on deep RGB–LiDAR fusion

Sep 2026 · International Conference on Intelligent Transportation Systems and Automation Control · Vol 14368, pp. 1436809 - 1436809-9 · 0 citations · 17 references
Engineering

TL;DR

A reliability-gated cross-modal fusion model, RG- CMF, is proposed for complex-scene 3D object recognition that integrates calibration-aware RGB–LiDAR alignment, dual- branch feature extraction, reliability-gated fusion, and geometry–semantic consistency optimization and suggests that explicit reliability modeling can improve the accuracy, robustness, and interpretability of RGB–LiDAR 3D object recognition under tested complex-scene conditions.

Abstract

Three-dimensional object recognition is an important perception task in autonomous driving, intelligent transportation, and mobile robotics, where models must identify object categories while estimating spatial location, size, and orientation. RGB cameras provide dense semantic and texture information, whereas LiDAR sensors provide accurate depth and geometric structure. However, camera-based methods are sensitive to illumination and depth ambiguity, while LiDAR- based methods may degrade under sparse point distribution, long-distance observation, and occlusion. Existing RGB– LiDAR fusion methods often rely on feature concatenation or fixed fusion weights, which insufficiently consider scene- dependent modal reliability. To address this issue, this paper proposes a reliability-gated cross-modal fusion model, RG- CMF, for complex-scene 3D object recognition. The method integrates calibration-aware RGB–LiDAR alignment, dual- branch feature extraction, reliability-gated fusion, and geometry–semantic consistency optimization. Experiments on the KITTI validation set show that RG-CMF achieves 84.7 ± 0.7, 57.1 ± 0.9, and 68.5 ± 0.9 moderate-level AP for cars, pedestrians, and cyclists, respectively. Compared with PV-RCNN, the mean translation error decreases from 0.351 ± 0.012 m to 0.319 ± 0.010 m. Ablation results further indicate that reliability-gated fusion contributes the largest performance gain, while consistency loss improves localization accuracy. The findings suggest that explicit reliability modeling can improve the accuracy, robustness, and interpretability of RGB–LiDAR 3D object recognition under tested complex-scene conditions.

View source

Similar papers

Open access Aug 2026

A Lightweight RGB-LiDAR Feature Recalibration Network for Large-Scale 3D Scene Understanding

Semantic segmentation of large-scale 3D point clouds is a fundamental task in robotic perception, semantic mapping, and urban scene understanding. Existing methods mainly rely on geometric information, which limits their ability to distinguish semantic categories with similar spatial structures. To address this issue,...

Weifeng Zhai, Zexi Tan · 0 citations
Oct 2026

Robust Multimodal Gated Fusion for 3-D Object Detection via Alignment and Denoising

Multimodal perception integrating light detection and ranging (LiDAR) and cameras has become a key paradigm for 3-D object detection, as it leverages both geometric structure and semantic information. However, in real-world autonomous driving scenarios, calibration errors, adverse weather, and sensor degradation can in...

Hui-Lin Huang, Yan Bai, Peng-Yuan Wang et al. · 0 citations
Conference Open access 2026

Automatic Driving 3D Object Detection Based on Lidar and Camera Information Feature Fusion

This review offers an all-round synthesis of how LiDAR and camera sensors are integrated for 3D target detection and aims to give a comprehensive theoretical explanation to scholars who have a preliminary understanding of the fusion of LiDAR and camera.

Yuan Liu · 0 citations
Conference Open access 2026

Multi-modal Feature Fusion in RGB-LiDAR SLAM

Simultaneous Localization and Mapping (SLAM) is a core technology for environmental perception and autonomous navigation in intelligent robots and autonomous driving. Monocular SLAM has inherent limitations and cannot adapt to complex real-world scenarios. This paper systematically reviews RGB-LiDAR multimodal fusion S...

Haoyang Jing · 0 citations
#artificial intelligence Preprint Sep 2026

SARFusion: Scene-Aware Routing Fusion for Robust Camera-LiDAR 3D Object Detection

Camera-LiDAR fusion has become a prevailing paradigm for 3D object detection in autonomous driving. However, existing fusion detectors often establish strong inter-modality dependencies by decoding object queries from tightly coupled multimodal representations. Under corrupted driving conditions, such dependencies make...

Yu-Ting Zhao, Zi-Yi Zheng, Shu-Xiao Li · 0 citations
Sep 2026

Toward High-Precision Target-Free LiDAR-Camera Extrinsic Calibration: Multi-Modal Geometric Edge Matching and Visibility-Guided Optimization

The rapid advancement of autonomous driving and embodied intelligence highlights the critical need for robust multi-sensor systems. In such systems, LiDAR and cameras are frequently paired—leveraging their complementary strengths: LiDAR excels at high-precision measurements, while cameras deliver rich texture informati...

Bao-Sheng Zhang, Lin Zhang, Sheng-Jie Zhao et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.