OccBEV-Oracle, a lightweight completion neck inserted between the BEV backbone and dense head of a CenterPoint-style detector, and a direct measurement of where the neck edits the BEV map, show that the aligned GT-ROM mask itself supplies a strong localization prior.
Abstract
LiDAR-based 3D object detection is sensitive to sparse point support and occlusion-induced incomplete bird’s-eye-view (BEV) representations, especially for pedestrians, cyclists, and distant objects. This paper asks how much accurate, target-aligned occlusion guidance can help BEV feature completion and why a geometrically estimated region-of-occlusion map (ROM) fails to reproduce that benefit. We introduce OccBEV-Oracle, a lightweight completion neck inserted between the BEV backbone and dense head of a CenterPoint-style detector. Given an occlusion mask and a density map, it selects low-density occluded tokens, aggregates visible-token context through density-weighted cross-attention, and applies spatially constrained weak-residual fusion. On the KITTI validation split, full-map completion at a strong residual coefficient reduces mean Moderate 3D AP_R40 from 59.42% to 56.65%, whereas oracle GT-ROM-guided masked completion raises mean Moderate and Hard AP_R40 to 64.31% and 60.63%. A matched-strength control shows that weakening the full-map residual recovers only part of this gain (60.97% mean Moderate), so spatial restriction contributes a further 3.34 points that residual strength alone cannot supply. The gains concentrate on Pedestrian and Cyclist and on partly occluded and middle/far-range objects. Raycasting-based estimated ROM variants remain below the baseline. Mismatch, shifted/shuffled, zero-context and mask-only controls, and a direct measurement of where the neck edits the BEV map, show that the aligned GT-ROM mask itself supplies a strong localization prior. A mask-target analysis on nuScenes confirms the failure mode is dataset-independent. OccBEV-Oracle is therefore an upper-bound analysis, not a deployable detector: accurate target-related occlusion localization remains the main bottleneck.
MVXCC-NET is presented, a cross-modal 3D detection network for occluded objects based on dual-path information complementation and regional weight modeling, which improves the utilization efficiency of fused features, allowing visual semantic information and spatial geometric information to support each other.
Jin Qi, Jian Wang· Journal of King Saud Univers...· 0 citations
Non-contact 3D measurements acquired by LiDAR, structured-light and depth-camera systems often contain large missing regions under occlusion, limited viewpoints, motion blur and sensor noise. These defects reduce the geometric fidelity of reconstructed shapes and directly affect downstream dimensional analysis, pose es...
Yu-Hao Yang, Gun Li, Jia-Cheng Luo et al.· Measurement science and tech...· 0 citations
Results validate the effectiveness of the proposed novel 3D object detection and tracking framework, termed ECF3DMOT, in advancing 3D object detection and tracking for autonomous driving.
Xiaojuan Peng, Fei Teng, Tiankai Chen et al.· International Journal of Mac...· 0 citations
Map-Det3D is an online multi-view 3D object detection model that brings detection directly into a 3D space reconstructed from RGB, suggesting that training reconstruction priors for detection is a practical route to stable metric 3D detection from monocular video.
Yung-Hsu Yang, Luigi Piccinelli, S. R. Bulò et al.· 1 citation
PH-PPC, a novel Point-Voxel based 3D object detection framework utilizing Height-Domain Attention and Point Cloud Pseudo-Completion, achieves competitive performance among voxel-based detectors and incorporates three key technical innovations to enhance robustness against occlusion.
Yuanlong Wang, Ze-Zheng Qing, Zijie Ji et al.· Machine Vision and Applicati...· 0 citations
This work introduces SAM-AD, a domain-specific pretraining strategy that fine-tunes SAM on autonomous-driving imagery to extract feature representations with rich semantic information, and develops the Depth-Guided Wavelet Attention (DGWA) module, which suppresses high-frequency sensor noise while preserving critical c...
Zi-Ying Song, Lin Liu, Hong-Yu Pan et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.