OccAnyScene is proposed, a pixel-frustum-centered Gaussian framework built upon a pretrained depth foundation model which employs Pixel-Aligned Frustum Feature Aggregation to construct a camera-aware frustum query for each feature pixel, and Frustum-Parameterized Gaussian Construction to decode each query into multiple Gaussians whose positions and sizes are constrained by the predicted pixel depth and corresponding frustum geometry.
Abstract
3D occupancy prediction is fundamental to scene understanding, yet existing 3D semantic occupancy methods are typically specialized to fixed scene types and occupancy protocols. We introduce Cross-Scene 3D Semantic Occupancy Prediction, a new task setting which requires a single model to handle heterogeneous indoor and outdoor scenes with varying cameras, spatial ranges, voxel specifications, and semantic taxonomies. This setting poses a fundamental challenge: achieving metric-consistent yet scene-adaptive image-to-3D lifting across varying camera configurations and scene scales. To address this challenge, we propose OccAnyScene, a pixel-frustum-centered Gaussian framework built upon a pretrained depth foundation model. Specifically, the framework employs Pixel-Aligned Frustum Feature Aggregation to construct a camera-aware frustum query for each feature pixel, and Frustum-Parameterized Gaussian Construction to decode each query into multiple Gaussians whose positions and sizes are constrained by the predicted pixel depth and corresponding frustum geometry. OccAnyScene sets new state-of-the-art results, achieving 59.92% mIoU on the indoor Occ-ScanNet and 23.06% mIoU on the outdoor SurroundOcc-nuScenes.
Semantic occupancy prediction gives an embodied agent a voxel-level account of where space is free, occupied, and semantically meaningful. In RGB-D pipelines such as EmbodiedScan, the image encoder is often left as a default module, even though its features are the visual evidence later sampled into the 3D grid. We stu...
OutLangSplat is presented which adapts language Gaussian representations to UAV outdoor scenes by improving feature representation and aggregation reliability, and is the first accessible dataset of open-vocabulary 3D scene understanding for UAV outdoor scenes.
Xiaosheng Yan, Hefeng Wu, Yanghui Xu et al.· 1 citation
3D occupancy prediction involves estimating the spatial structure and occupancy state of each voxel within a scene. Vision-centric approaches have attracted increasing attention for their cost-effectiveness and ease of deployment. However, a key challenge remains in accurately inferring the 3D scene structure from plan...
Ke-Qiu Wang, Tian-Yu Shen, Si-Han Chen et al.· IEEE Transactions on Automat...· 0 citations
Camera-based autonomous driving perception requires a shared representation that preserves metric 3D structure across synchronized multi-camera streams. However, existing image-based frameworks often rely on backbones pretrained for semantic recognition, and introduce 3D geometry through downstream task-specific module...
Longfei Xu, Xiao-Hui Wang, Ze-Hao Huang et al.· 2 citations
UniQuery4R is presented, a query-conditioned framework that encodes a multi-frame clip once and selects the source view, target view, and continuous source-image coordinate only at decoding time via source-to-target cross-attention, and introduces a direction-magnitude parameterization of scene flow with separate super...
Tiancheng Chen, Sheng Tang, Wenhua Jin et al.· 0 citations
Instance Image-Goal Navigation (IIN) asks an agent to locate the specific object instance shown in a goal image. Existing 3D Gaussian Splatting (3DGS) based methods rely on pose-centric search—sampling many viewpoints, rendering them, and comparing against the goal—which is inefficient in continuous 6-DoF space. We ins...
Yijie Deng, Shuaihang Yuan, Geeta Chandra Raju Bethala et al.· IEEE Robotics and Automation...· 3 citations· ⚡1
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.