EO-VGGT is presented, a framework that adapts a frozen perspective-driven model to orbital observations via explicit physical geometry embedding via explicit physical geometry embedding for robust feed-forward satellite 3D reconstruction.
Abstract
In the era of satellite constellations, multi-view optical satellite imagery is pivotal for Earth Observation (EO) and high-quality Digital Surface Model (DSM) reconstruction. Although feed-forward 3D foundation models have transformed computer vision, their deployment in satellite remote sensing is inherently constrained by the structural discrepancy between implicit perspective assumptions and explicit orbital pushbroom geometry. This geometric incongruity is further compounded by pronounced view-set heterogeneity. We present EO-VGGT, a framework that adapts a frozen perspective-driven model to orbital observations via explicit physical geometry embedding.First, the Geometry-Correlation Constrained Selection (GCCS) strategy prunes sub-optimal observations by balancing geometric diversity and radiometric consistency to optimize the input sequence. Second, a Sensor-Ray Encoder (SRE) parameterizes pixel-level pushbroom lines of sight derived from the Rational Function Model (RFM) into high-dimensional space-geometric tokens, reconciling the mathematical discrepancy between central projection and orbital kinematics. Third, a lightweight Ray-Pointing-Aware Adapter (RPAA) employs gated residual blocks to integrate these tokens directly into the frozen transformer backbone. Our findings underscore that integrating explicit physical geometry with optimized view selection is essential for robust feed-forward satellite 3D reconstruction.
Abstract. Satellite imagery offers a distinct advantage in Earth observation by providing expansive coverage and enabling the monitoring of inaccessible regions without physical on-site intervention, serving as a significantly more cost-effective and scalable alternative to traditional aerial or ground-based surveys. The task of 3D reconstruction from multi-view satellite images has therefore been a pivotal point of research at the intersection of photogrammetry and remote sensing. Recently, novel-view synthesis techniques such as Neural Radiance Fields (NeRF) and 3D Gaussian Splatting (3DGS) have accelerated the accuracy and speed of topographic modeling. Among these, Earth Observation Gaussian Splatting (EOGS) has emerged as a state-of-the-art approach by adapting 3DGS to handle the unique geometric and radiometric characteristics of satellite data, including Rational Polynomial Coefficients (RPCs) and varying solar conditions. However, the standard EOGS pipeline relies on stochastic initialization, where Gaussians are distributed uniformly within a volumetric bounding box, leading to high computational overhead and dependency on aggressive pruning that can inadvertently remove critical geometric features, particularly in areas with complex urban structures. To address these limitations, we propose Bundle-Adjusted Initialization for Earth Observation Gaussian Splatting, which leverages sparse point clouds from bundle adjustment as geometric priors for Gaussian initialization. Combined with an adaptive densification strategy, our method achieves faster convergence and improved DSM accuracy on the DFC2019 dataset compared to the EOGS baseline.
Jiyong Kim, Shuang Song, Rongjun Qin· The International Archives o...· 0 citations
Abstract. The emergence of satellite constellations enables near-synchronous multi-view optical imaging, offering new opportunities for large-scale 3D city modeling. Yet a practically promising configuration, in which a primary near-nadir view is complemented by multiple oblique side-looking viewpoints, remains under-examined. This study develops a controlled semi-simulation framework to analyze how multi-view imaging geometry affects the recoverability of urban 3D structures. Under idealized conditions with imaging perturbations removed, e.g., radiometric, illumination, and sensor model errors, the experiments focus on three practical factors: the number of side-looking views, view obliqueness, and the constellation’s azimuthal orientation relative to the scene. With parameter sweep analysis, it reveals an asymmetric U-shaped trend between reconstruction performance and both the view count and the obliqueness: moderate angular diversity markedly strengthens urban scene recoverability. In contrast, large obliqueness reduces inter-view overlap and destabilizes matching, while excessive redundancy introduces consistency issues that ultimately degrade reconstruction performance. Furthermore, the results shows that geometric accuracy, completeness, and texture appearance each peak at different parameter combinations, revealing intrinsic trade-offs in multi-view urban reconstruction, as different evaluation criteria favor distinct optimal configurations. The study provides practical guidance for the geometric design and mission planning of multi-satellite constellations aimed at improving satellite-based 3D modeling in urban areas.
Xu Cheng, Xianfeng Huang, Y. Pi et al.· ISPRS Annals of the Photogra...· 0 citations
Overall, this study systematically quantifies the performance of 3D VFMs in satellite image-based 3D reconstruction, confirming their strong potential for high-resolution satellite applications and providing valuable insights for enhancing model robustness and generalization across complex urban and low-resolution environments.
Liupeng Su, Yuhao Ye, Han Hu et al.· ISPRS Annals of the Photogra...· 0 citations
High-precision 3D reconstruction of on-orbit non-cooperative targets is essential for space situational awareness. However, extreme space environments induce severe imaging degradations, including high-dynamic-range (HDR) illumination, rapid motion blur, and platform jitter. Traditional 3D Gaussian Splatting (3DGS) conflates these optical distortions with geometric optimization, leading to pathological structural inflation and the loss of thin appendages like solar panels. To overcome this, we propose OrbitGS, a physically decoupled 3DGS framework. OrbitGS integrates physical imaging priors via a Kinematics-Driven Degradation Synthesizer (KDDS) to deterministically extract view-specific degradation kernels. Furthermore, a blur-decoupled rendering strategy with intensity-aware weighting mitigates HDR variations, while a semantic-aware densification scheme mathematically penalizes abnormal primitive expansion. Evaluations on the SPE3R dataset demonstrate that OrbitGS effectively disentangles optical degradations from the geometric representation. Quantitatively, evaluated across seven space targets under moderate (200-view) and extreme sparse (50-view) settings, our framework achieves state-of-the-art robustness against extreme degradations. Notably, it avoids the catastrophic structural blow-ups observed in baseline methods, yielding an average geometric F1-score of 0.80 and a Chamfer Distance of 0.818, alongside a rendering Structural Similarity Index (SSIM) of 0.82 and a Learned Perceptual Image Patch Similarity (LPIPS) of 0.16. By preserving delicate structures under severe degradation, OrbitGS provides a robust, high-fidelity 3D reconstruction solution for complex orbital environments.
Ligang Li, Ziyan Qin, Fan Zhang et al.· Remote Sensing· 0 citations
Synthetic aperture radar (SAR) image generation can mitigate data scarcity, but controllablegeneration under sparse observation angles remains difficult. Recent SAR generative studies im-prove texture realism, yet explicit geometry-aware control is still limited. This paper studiesthe focused and verifiable setting of intermediate-azimuth completion: 3D-model-derived geo-metric priors guide a diffusion model to synthesize the views missing from sparse-angle trainingdata. GeoDiff-SAR constructs a lightweight multi-bounce ray-tracing prior, encodes the result-ing point cloud, and fuses it with text conditioning while adapting Stable Diffusion 3.5 Mediumthrough low-rank adaptation. On a real four-category aircraft dataset, GeoDiff-SAR reaches anSSIM of 0.812 and azimuth consistency of 0.940, compared with 0.738 and 0.782 for the text-conditioned SD3.5 Medium baseline. The same sparse-angle protocol on five MSTAR vehicleclasses yields an SSIM of 0.878 and azimuth consistency of 0.917. These results support theconclusion that a lightweight 3D geometric prior improves viewpoint adherence for controllableSAR generation; it is intended as generation guidance rather than high-fidelity electromagneticreconstruction.
Fan Zhang, Xuanting Wu, Fei Ma et al.· 0 citations
A platform-aware benchmark framework that jointly records visual fidelity, computational cost, metric geometry, product utility, failure behavior, and reproducibility metadata for UAV/aerial, satellite, and hybrid settings is proposed.
Wen-Xuan Fan, Bo Wang, Junqiang Ye et al.· Remote Sensing· 0 citations