Skip to content
Preprint

OccAnyScene: Towards Unified Indoor-Outdoor 3D Occupancy Prediction

Aug 2026 · 0 citations · 38 references
Computer Science

TL;DR

OccAnyScene is proposed, a pixel-frustum-centered Gaussian framework built upon a pretrained depth foundation model which employs Pixel-Aligned Frustum Feature Aggregation to construct a camera-aware frustum query for each feature pixel, and Frustum-Parameterized Gaussian Construction to decode each query into multiple Gaussians whose positions and sizes are constrained by the predicted pixel depth and corresponding frustum geometry.

Abstract

3D occupancy prediction is fundamental to scene understanding, yet existing 3D semantic occupancy methods are typically specialized to fixed scene types and occupancy protocols. We introduce Cross-Scene 3D Semantic Occupancy Prediction, a new task setting which requires a single model to handle heterogeneous indoor and outdoor scenes with varying cameras, spatial ranges, voxel specifications, and semantic taxonomies. This setting poses a fundamental challenge: achieving metric-consistent yet scene-adaptive image-to-3D lifting across varying camera configurations and scene scales. To address this challenge, we propose OccAnyScene, a pixel-frustum-centered Gaussian framework built upon a pretrained depth foundation model. Specifically, the framework employs Pixel-Aligned Frustum Feature Aggregation to construct a camera-aware frustum query for each feature pixel, and Frustum-Parameterized Gaussian Construction to decode each query into multiple Gaussians whose positions and sizes are constrained by the predicted pixel depth and corresponding frustum geometry. OccAnyScene sets new state-of-the-art results, achieving 59.92% mIoU on the indoor Occ-ScanNet and 23.06% mIoU on the outdoor SurroundOcc-nuScenes.

View source

Similar papers

Preprint Sep 2026

Exploring 2D backbone effects for indoor semantic occupancy prediction

Semantic occupancy prediction gives an embodied agent a voxel-level account of where space is free, occupied, and semantically meaningful. In RGB-D pipelines such as EmbodiedScan, the image encoder is often left as a default module, even though its features are the visual evidence later sampled into the 3D grid. We stu...

Shizhang Fanga, Wanling Yea, Qi Zheng · 0 citations
Preprint Aug 2026

OutLangSplat: 3D Language Gaussian Splatting for UAV Outdoor Scenes

OutLangSplat is presented which adapts language Gaussian representations to UAV outdoor scenes by improving feature representation and aggregation reliability, and is the first accessible dataset of open-vocabulary 3D scene understanding for UAV outdoor scenes.

Xiaosheng Yan, Hefeng Wu, Yanghui Xu et al. · 1 citation
2026

GD-OCC: Geometry-Driven 3D Occupancy Prediction With Integrated View Transformation and Sparse Representation

3D occupancy prediction involves estimating the spatial structure and occupancy state of each voxel within a scene. Vision-centric approaches have attracted increasing attention for their cost-effectiveness and ease of deployment. However, a key challenge remains in accurately inferring the 3D scene structure from plan...

Ke-Qiu Wang, Tian-Yu Shen, Si-Han Chen et al. · 0 citations
Preprint Aug 2026

Geometry-Grounded Unified 3D Perception for Autonomous Driving

Camera-based autonomous driving perception requires a shared representation that preserves metric 3D structure across synchronized multi-camera streams. However, existing image-based frameworks often rely on backbones pretrained for semantic recognition, and introduce 3D geometry through downstream task-specific module...

Longfei Xu, Xiao-Hui Wang, Ze-Hao Huang et al. · 2 citations
Preprint Aug 2026

UniQuery4R: Unified 4D Scene Reconstruction from a Single Query

UniQuery4R is presented, a query-conditioned framework that encodes a multi-frame clip once and selects the source view, target view, and continuous source-image coordinate only at decoding time via source-to-target cross-attention, and introduces a direction-magnitude parameterization of scene flow with separate super...

Tiancheng Chen, Sheng Tang, Wenhua Jin et al. · 0 citations
Open access Jun 2025

Hierarchical Scoring With 3D Gaussian Splatting for Instance Image-Goal Navigation

Instance Image-Goal Navigation (IIN) asks an agent to locate the specific object instance shown in a goal image. Existing 3D Gaussian Splatting (3DGS) based methods rely on pose-centric search—sampling many viewpoints, rendering them, and comparing against the goal—which is inefficient in continuous 6-DoF space. We ins...

Yijie Deng, Shuaihang Yuan, Geeta Chandra Raju Bethala et al. · 3 citations · ⚡1

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.