Skip to content
Open access

Decoupling Mask Quality From Completion Design: A Diagnostic Framework for Occlusion-Aware LiDAR 3-D Object Detection

2026 · IEEE Access · Vol 14, pp. 128636-128657 · 0 citations · 59 references

TL;DR

OccBEV-Oracle, a lightweight completion neck inserted between the BEV backbone and dense head of a CenterPoint-style detector, and a direct measurement of where the neck edits the BEV map, show that the aligned GT-ROM mask itself supplies a strong localization prior.

Abstract

LiDAR-based 3D object detection is sensitive to sparse point support and occlusion-induced incomplete bird’s-eye-view (BEV) representations, especially for pedestrians, cyclists, and distant objects. This paper asks how much accurate, target-aligned occlusion guidance can help BEV feature completion and why a geometrically estimated region-of-occlusion map (ROM) fails to reproduce that benefit. We introduce OccBEV-Oracle, a lightweight completion neck inserted between the BEV backbone and dense head of a CenterPoint-style detector. Given an occlusion mask and a density map, it selects low-density occluded tokens, aggregates visible-token context through density-weighted cross-attention, and applies spatially constrained weak-residual fusion. On the KITTI validation split, full-map completion at a strong residual coefficient reduces mean Moderate 3D AP_R40 from 59.42% to 56.65%, whereas oracle GT-ROM-guided masked completion raises mean Moderate and Hard AP_R40 to 64.31% and 60.63%. A matched-strength control shows that weakening the full-map residual recovers only part of this gain (60.97% mean Moderate), so spatial restriction contributes a further 3.34 points that residual strength alone cannot supply. The gains concentrate on Pedestrian and Cyclist and on partly occluded and middle/far-range objects. Raycasting-based estimated ROM variants remain below the baseline. Mismatch, shifted/shuffled, zero-context and mask-only controls, and a direct measurement of where the neck edits the BEV map, show that the aligned GT-ROM mask itself supplies a strong localization prior. A mask-target analysis on nuScenes confirms the failure mode is dataset-independent. OccBEV-Oracle is therefore an upper-bound analysis, not a deployable detector: accurate target-related occlusion localization remains the main bottleneck.

Read PDF

Similar papers

Open access Aug 2026

MVXCC-NET: Cross-modal 3D detection of occluded objects based on dual-path information complementation and regional weight modeling

MVXCC-NET is presented, a cross-modal 3D detection network for occluded objects based on dual-path information complementation and regional weight modeling, which improves the utilization efficiency of fused features, allowing visual semantic information and spatial geometric information to support each other.

Jin Qi, Jian Wang · 0 citations
Sep 2026

Cross-modal point cloud completion for robust non-contact 3D measurement under occlusion

Non-contact 3D measurements acquired by LiDAR, structured-light and depth-camera systems often contain large missing regions under occlusion, limited viewpoints, motion blur and sensor noise. These defects reduce the geometric fidelity of reconstructed shapes and directly affect downstream dimensional analysis, pose es...

Yu-Hao Yang, Gun Li, Jia-Cheng Luo et al. · 0 citations
Aug 2026

Ecf3dmot: enhanced centerpoint framework for 3D object detection and tracking with LiDAR

Results validate the effectiveness of the proposed novel 3D object detection and tracking framework, termed ECF3DMOT, in advancing 3D object detection and tracking for autonomous driving.

Xiaojuan Peng, Fei Teng, Tiankai Chen et al. · 0 citations
Preprint Aug 2026

Map-Det3D: Metric Feed-Forward 3D Reconstruction Prior for Multi-view 3D Object Detection from Streaming Inputs

Map-Det3D is an online multi-view 3D object detection model that brings detection directly into a 3D space reconstructed from RGB, suggesting that training reconstruction priors for detection is a practical route to stable metric 3D detection from monocular video.

Yung-Hsu Yang, Luigi Piccinelli, S. R. Bulò et al. · 1 citation
Sep 2026

Voxel based 3D object detection using SE attention mechanism in height domain and point cloud pseudo-completion

PH-PPC, a novel Point-Voxel based 3D object detection framework utilizing Height-Domain Attention and Point Cloud Pseudo-Completion, achieves competitive performance among voxel-based detectors and incorporates three key technical innovations to enhance robustness against occlusion.

Yuanlong Wang, Ze-Zheng Qing, Zijie Ji et al. · 0 citations
Preprint Sep 2026

Towards robust multimodal 3D object detection via visual foundation models

This work introduces SAM-AD, a domain-specific pretraining strategy that fine-tunes SAM on autonomous-driving imagery to extract feature representations with rich semantic information, and develops the Depth-Guided Wavelet Attention (DGWA) module, which suppresses high-frequency sensor noise while preserving critical c...

Zi-Ying Song, Lin Liu, Hong-Yu Pan et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.