Sep 2026· IEEE transactions on intelligent transportation systems (Print)· Vol 27, pp. 10543-10553· 0 citations· 74 references
Abstract
Three-dimensional object detection is a key component of the perception module in autonomous driving systems. Compared to camera images, LiDAR point clouds provide richer spatial information, such as detailed structural and geometric cues of objects. However, existing 3D object detection methods face two major challenges: 1) loss of critical information due to downsampling or convolution operations, leading to missed detections, and 2) insufficient exploitation of contextual information around objects, resulting in inaccurate bounding box regression. These problems are particularly severe when detecting sparse point cloud objects that are distant or occluded. To address these issues, we propose two novel modules applicable to both point-based and voxel-based detection frameworks: the Class-Enhanced Multi-Sampling (CEMS) module and the Multi-Level Graph Attention (MLGA) module. The CEMS module employs an iterative class-aware sampling and semantic interpolation strategy to achieve more precise downsampling of critical information, while the MLGA module integrates both intra-layer and cross-layer graph attention to capture local geometric structures around objects as well as inter-object relationships, thereby enhancing the representation of key features. Comprehensive evaluations on multiple datasets, including ONCE, Waymo, and nuScenes, demonstrate that the proposed modules can consistently improve the detection performance across various network architectures while maintaining high inference efficiency.
In the realm of autonomous driving, 3D object detection based on LiDAR point clouds has emerged as a pivotal technology. To enhance the accuracy and efficiency of 3D object detection, this paper introduces Spatial-Channel Attention Guided with Gumbel Subset Sampling and Context Fusion RCNN (SCAGCF-RCNN), a two-stage...
Hong-Xu Li, Shu-Yi Zhou, Yi-Tao Lu et al.· Engineering Research Express· 0 citations
A aggregated Euclidean distance weighted box fusion method, which aggregates complementary information from multiple candidate boxes during post-processing to improve bounding-box selection and localization accuracy, and a hybrid deformable half-conv (HDHC) module that jointly enhances global and local feature represen...
Di Tian, Jia-Wei Wang, Jia-Bo Li et al.· Measurement science and tech...· 0 citations
MVXCC-NET is presented, a cross-modal 3D detection network for occluded objects based on dual-path information complementation and regional weight modeling, which improves the utilization efficiency of fused features, allowing visual semantic information and spatial geometric information to support each other.
Jin Qi, Jian Wang· Journal of King Saud Univers...· 0 citations
This work proposes a cascade optimization framework that systematically enhances feature representation and refines multimodal fusion, and introduces the Multi-Scale Contextual Fusion Module (MSCF) to reduce alignment bias.
This work introduces SAM-AD, a domain-specific pretraining strategy that fine-tunes SAM on autonomous-driving imagery to extract feature representations with rich semantic information, and develops the Depth-Guided Wavelet Attention (DGWA) module, which suppresses high-frequency sensor noise while preserving critical c...
Zi-Ying Song, Lin Liu, Hong-Yu Pan et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.