Multi-modal Feature Fusion in RGB-LiDAR SLAM
Abstract
Simultaneous Localization and Mapping (SLAM) is a core technology for environmental perception and autonomous navigation in intelligent robots and autonomous driving. Monocular SLAM has inherent limitations and cannot adapt to complex real-world scenarios. This paper systematically reviews RGB-LiDAR multimodal fusion SLAM, verifying the robustness deficiency of RGB visual SLAM and semantic information loss of LiDAR SLAM via literature analysis and experimental quantification to demonstrate fusion necessity. It sorts out three fusion paradigms (data- level, feature-level, decision-level), analyzes the evolution of fusion technology in the whole SLAM process from traditional methods to deep learning-driven ones, and verifies fusion effects in indoor dynamic and outdoor unstructured scenarios. Furthermore, it quantifies two core challenges including cross-modal registration error and edge-side lightweight contradiction, and proposes three future directions including self-supervised cross-modal registration, lightweight transformer-based fusion, and semantic-geometric joint constraint. Experimental results show RGB-LiDAR fusion improves SLAM localization accuracy by over 60% and robustness by about 300% in complex scenarios, achieving geometric- semantic integrated mapping. This work provides theoretical references for the engineering implementation of multimodal fusion SLAM.