Skip to content
Conference

FP-RCNN: a feature scale enhanced multimodal fusion method for object detection

Jul 2026 · International Conference on Sensor Technology and Information Engineering · Vol 14258, pp. 142580N - 142580N-7 · 0 citations · 7 references
Engineering

Abstract

Aiming at low object detection accuracy from feature scale mismatch and local feature loss in RGB-Lidar multi-modal fusion for unmanned scenes, we propose FP-RCNN, a 3D object detection method based on Feature Fusion Pyramid Attention (FFPA). To solve scale mismatch, a fusion strategy of multi-scale feature matching and double self-attention superimposition is introduced in feature recognition: multi-scale feature maps are obtained via a feature pyramid network, with important regions recalibrated by double self-attention. To address local feature loss, the point cloud segmentation network is optimized in 3D instance segmentation; input point sets connect local and global features via shared MLP and NetVLAD for disorder invariance and higher fusion accuracy. KITTI dataset experiments show FP-RCNN significantly improves detection accuracy, especially in challenging scenes: 78.07% (difficult cars), 50.09%/45.77% (medium/difficult pedestrians), and 60.91% (difficult cyclists), surpassing selected comparison algorithms. This research advances 3D object detection in autonomous driving and provides a robust solution for complex environment navigation.

View source