Skip to content

EDAFusion: LiDAR-Guided Multilevel Depth Enhancement and Dynamic Scale Attention Fusion for Multimodal 3-D Object Detection

2026 · IEEE Transactions on Instrumentation and Measurement · Vol 75, pp. 2517814-2517814 · 0 citations · 40 references

Abstract

Three-dimensional object detection plays a critical role in intelligent robotics and autonomous driving, where accurate and robust perception remains challenging under multimodal fusion settings. Existing bird’s-eye-view (BEV)-based multimodal methods still suffer from unreliable camera depth estimation, insufficient adaptation to object scale variations, and temporal misalignment caused by dynamic scenes. To address these challenges, this article proposes EDAFusion, a unified multimodal 3-D object detection framework with enhanced BEV generation and fusion. In particular, a LiDAR-guided multilevel depth enhancement (LMDE) strategy is introduced to improve the robustness of camera BEV representation under challenging illumination conditions; a dynamic scale attention fusion (DSAF) module is designed to enhance cross-modal interaction across objects of different scales; and a motion-guided temporal aggregation (MTA) method is developed to improve temporal consistency under object motion and partial occlusion. Experiments on the nuScenes test set demonstrate that EDAFusion achieves 72.1 % mean average precision (mAP) and 74.5 % nuScenes detection score (NDS), outperforming BEVFusion by 1.9 % mAP and 1.6 % NDS. These results demonstrate the effectiveness of the proposed framework for reliable multimodal 3-D perception in intelligent measurement and sensing systems.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.