This work reformulates DragOSM into DragRoof, an ODE-based framework inspired by human annotation behavior, and introduces the Oblique City dataset (ObliCity), the first large-scale benchmark that integrates high-resolution UAV imagery and globally distributed satellite data, covering diverse city morphologies and camera perspectives.
Abstract
Oblique-view urban remote sensing imagery inevitably exhibits geometric projection displacements between building roofs and footprints, leading to significant distortions in spatial structure. Existing approaches either ignore these deformations or handle them implicitly within segmentation-based frameworks, where progress is dominated by general segmentation advances rather than improvements in geometric correction. In this work, we explicitly define roof-to-footprint offset vector (RFOV) extraction as an independent learning task that decouples geometric alignment from semantic segmentation. To support this task, we introduce the Oblique City dataset (ObliCity), the first large-scale benchmark that integrates high-resolution UAV imagery and globally distributed satellite data, covering diverse city morphologies and camera perspectives. Methodologically, we reformulate DragOSM into DragRoof, an ODE-based framework inspired by human annotation behavior. By simulating the continuous process of dragging roofs toward their footprints, DragRoof learns deterministic, geometry-consistent offset fields and adaptively determines convergence through an end token. Extensive experiments on ObliCity demonstrate that DragRoof achieves state-of-the-art RFOV extraction performance, requiring fewer inference steps while delivering superior directional and length accuracy. Our dataset and model establish a principled foundation for studying projection displacement correction in oblique remote sensing imagery. The source code and dataset will be avaliable at https://github.com/likaiucas/DragRoof.
Extracting building footprints from aerial or satellite imagery remains a significant challenge, particularly in maintaining the geometric regularity of man-made structures. While polygon-based methods offer vectorized representations superior to pixel-based approaches, they often struggle with corner ambiguity and fai...
Wen-Jie Zhao, Xue-Jing Xie, Ze Meng et al.· IEEE Transactions on Geoscie...· 0 citations
Accurate semantic segmentation of large-scale outdoor LiDAR point clouds remains a challenging endeavor, primarily due to ambiguous class transitions at object interfaces, non-uniform sampling density across the surveyed area, and shared geometric signatures among distinct object categories. This paper proposes GFE-Net...
Coastal environments contain rich, largely unexploited geometric structure capable of providing globally referenced localization cues. In this work, we present two complementary localization frameworks that exploit shoreline and water-surface geometry for GPS-denied autonomous surface vessel localization. The first fra...
Large-scale 3D surface reconstruction from aerial imagery is fundamental to geospatial mapping and urban modeling. Recent advances in 3D Gaussian Splatting (3DGS) have demonstrated considerable potential for this task. However, existing methods still face three major challenges in large and complex scenes: scene partit...
Bo-Cheng Li, Wen-Juan Zhang, Jie-Pan-Dong-Xu Han et al.· 0 citations
Ground image localization with respect to satellite imagery is a key enabler for metrically-accurate, geo-localized 3D scene reconstruction from unconstrained image collections. Existing cross-view localization methods have strict requirements such as panoramic imagery or known initial locations, limiting their applica...
A. Daruna, Ben Southall, Niluthpol Chowdhury Mithun et al.· 0 citations
GeoScene is a geospatially guided framework that jointly uses satellite imagery and structured OpenStreetMap cues as soft priors for 3D semantic scene completion and consistently improves both geometric and semantic completion under the geospatial-prior-assisted setting.
Meng Wang, Shougao Zhang, Wenzhe He et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.