Skip to content
Preprint

Seed2GS: Camera-Free, Training-Free Object Extraction from 3D Gaussian Scenes via a Single Reference-View Grounding

Aug 2026 · 0 citations · 19 references
Computer Science

TL;DR

Seed2GS is presented, which achieves the highest reported LERF-MASK accuracy without original reconstruction cameras or scene-specific representation training, and its key insight is to separate target identity from 3D coverage.

Abstract

Extracting a target object from a pre-built 3D Gaussian Splatting (3DGS) scene enables interactive 3D editing. Existing methods either train for tens of minutes per scene, sacrifice accuracy, or require original reconstruction cameras that pre-built assets may not include. We present Seed2GS, which achieves the highest reported LERF-MASK accuracy without original reconstruction cameras or scene-specific representation training. Its key insight is to separate target identity from 3D coverage. QD-SAM3 selects one reliable reference mask from several open-vocabulary candidates, fixing identity once. Seed lift and visibility-adaptive virtual orbits then expose the object from new viewpoints, while tracking propagates the seed without repeated detection. Because the scene remains frozen, these masks supervise only one temporary foreground logit per Gaussian. On LERF-MASK, Seed2GS reaches 92.1% mean intersection over union (mIoU) with a measured compute-only latency of 9.3 seconds, 3.7 points above the strongest scene-trained baseline and 7.6 points above the closest camera-free baseline. With one fixed test reference per scene, the complete pipeline retains 91.1% mIoU; replacing its predicted seed with a ground-truth mask improves mIoU by only 0.72 points. On 3D-OVS, Seed2GS reaches 95.7% mIoU.

View source

Similar papers

Preprint Aug 2026

Map-Det3D: Metric Feed-Forward 3D Reconstruction Prior for Multi-view 3D Object Detection from Streaming Inputs

Map-Det3D is an online multi-view 3D object detection model that brings detection directly into a 3D space reconstructed from RGB, suggesting that training reconstruction priors for detection is a practical route to stable metric 3D detection from monocular video.

Yung-Hsu Yang, Luigi Piccinelli, S. R. Bulò et al. · 1 citation
Open access 2026

Decoupling Mask Quality From Completion Design: A Diagnostic Framework for Occlusion-Aware LiDAR 3-D Object Detection

OccBEV-Oracle, a lightweight completion neck inserted between the BEV backbone and dense head of a CenterPoint-style detector, and a direct measurement of where the neck edits the BEV map, show that the aligned GT-ROM mask itself supplies a strong localization prior.

Jun Wang, Quanxin Zheng, Jian-Ping Yu · 0 citations
Preprint Aug 2026

NCGR: Noise-Conditional Gated Rectification for Camera Extrinsic Perturbations in BEV 3D Object Detection

Camera-based bird's-eye-view (BEV) 3D detection typically assumes accurate and fixed camera extrinsics. In detectors using spatial cross-attention (SCA), extrinsic perturbations displace the image-plane projections of BEV reference points, causing queries to sample features from incorrect regions and degrading detectio...

Wenbin Pan, Wanhao Liu, Liwei Luo et al. · 0 citations
Preprint Aug 2026

CDSeg: A Renderable Gaussian Carrier for Image-to-3D Label Transfer

Modern image models provide strong cues about \emph{what} should be segmented in each view, but their masks do not by themselves determine \emph{where} those labels should persist in 3D. We present Cross-Domain Segmentation via Gaussian Splatting (CDSeg), a label-transfer interface that requires no task-specific 3D seg...

Wen-Tao Sun, Yi-Ping Chen, Zheng-Sen Xu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.