Skip to content

GD-OCC: Geometry-Driven 3D Occupancy Prediction With Integrated View Transformation and Sparse Representation

2026 · IEEE Transactions on Automation Science and Engineering · Vol 23, pp. 16512-16524 · 0 citations · 53 references

Abstract

3D occupancy prediction involves estimating the spatial structure and occupancy state of each voxel within a scene. Vision-centric approaches have attracted increasing attention for their cost-effectiveness and ease of deployment. However, a key challenge remains in accurately inferring the 3D scene structure from planar view information, primarily due to the difficulties in precise depth modeling during view transformation and the neglect of geometric sparsity in dense occupancy representations. To overcome the above challenges, a novel geometry-driven framework for occupancy prediction (GD-OCC) is proposed. A geometry-driven view transformation module is introduced, which integrates geometric cues into both explicit transformation and voxel-based implicit view transformation, thereby enhancing the quality of occupancy features. This design enables geometry-consistent 3D structure recovery of a scene from multi-view 2D images. Furthermore, a geometry-guided sparse voxel proposal strategy is presented to directly retain potentially non-empty voxels, facilitating a sparse occupancy representation with improved geometric fidelity. The effectiveness of our proposed GD-OCC is demonstrated through extensive experiments on the Occ3D-nuScenes dataset, achieving a RayIoU of 37.9% and outperforming existing methods, with particularly strong performance in long-tail scenarios. The code and models are available on github. Note to Practitioners—Accurate 3D scene perception is critical for autonomous vehicles, mobile robots, and other automation systems. Traditional camera-based methods are often affected by depth errors during view transformation and redundant voxels, limiting reliability and robustness. GD-OCC integrates geometric information into view transformation and focuses on potentially non-empty voxels, effectively improving perception accuracy, reducing unnecessary computation, and enhancing safe navigation and environment understanding. The method relies on multi-view images but can be easily extended to multi-sensor systems. Other potential applications include warehouse automation, robotic inspection, infrastructure monitoring, construction site management, and any automation tasks requiring reliable and efficient 3D environment modeling.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.