SVI2LoD3: Agent-Driven Reconstruction of LoD3 Façade Openings in Semantic 3D City Models from Volunteered Street View Imagery using Large Language and Visual Models
Aug 2026· ISPRS Annals of the Photogrammetry, Remote Sensing and Spatial Information Sciences· 1 citation· 9 references
Computer Science
TL;DR
An end-to-end, agent-driven pipeline for the LoD3 reconstruction of façade openings in 3D city models, producing directly usable CityGML-conform outputs and a novel evaluation metric for façade reconstruction, termed Facade Feature Distance.
Abstract
This paper presents an end-to-end, agent-driven pipeline for the LoD3 reconstruction of façade openings in 3D city models, producing directly usable CityGML-conform outputs. In contrast to existing approaches that rely on supervised semantic segmentation and therefore require large amounts of manually annotated training data, the proposed method employs a zero-shot segmentation strategy. This substantially reduces the annotation effort while still achieving strong performance in our benchmark on the eTRIMS dataset. A further key contribution is the enforcement of correct partonomic hierarchies, thereby producing CityGMLconform LoD3 building models. Beyond the reconstruction pipeline itself, this work also introduces a novel evaluation metric for façade reconstruction, termed Facade Feature Distance (FFD). Unlike conventional metrics such as mIoU or FRDS, which assess similarity primarily through pixel-wise overlap, FFD measures distance in a high-level feature space derived from a vision transformer. In doing so, it captures both semantic correctness and architectural layout, providing a more suitable assessment of façade reconstruction quality. The proposed pipeline and evaluation strategy together offer a practical and scalable contribution toward the automated generation and analysis of semantically enriched 3D city models. The developed code is published at: https://github.com/hcu-cml/citydb-SVI2LoD3-ai.
CityGML Level of Detail 3 (LOD3) building model generation requires accurate façade representation. Existing workflows often rely on LiDAR, dense multi-view imagery, or pre-existing lower-LOD models, limiting their applicability in data-scarce environments. Additionally, most methods are developed on buildings with sim...
Demi Julianna L. Gentiles, Khalil C. Torneros, J. Casisirano et al.· The International Archives o...· 0 citations
Regular road inspection is essential for maintaining pavement service life and ensuring traffic safety. Conventional approaches require specialized survey vehicles and sensor payloads, while damage assessment relies predominantly on supervised computer vision models trained on large annotated datasets. Recent advances...
B. Ouzougarh, Jannik Matijevic, Huynh Duc An Son Nguyen et al.· ISPRS Annals of the Photogra...· 0 citations
The results show that distributionally-regularized joint embedding architectures can be successful on challenging city-scale 3D scenes, and that transfer improves when self-supervision is designed for the capture geometry and spatial context of this domain while also revealing the limits of this specialization.
A. Rusnak, S. Kovalenko, Jingru Wang et al.· 0 citations
Accurate 3D semantic perception is critical for safe autonomous navigation. However, supervised LiDAR segmentation remains tied to closed taxonomies and to the cost of point-wise manual annotation. Open-vocabulary methods avoid that cost by projecting the output of 2D vision-language models onto LiDAR and distilling it...
Cigdem Kokenoz, Amir Salarpour, Alkim Domeke et al.· 0 citations
Urban applications, such as facade reconstruction, digital twins, building analysis, and semantic 3-D city modeling benefit from reliable and geometrically consistent representations of window and door (WD) elements. Many image-based facade reconstruction pipelines first detect WD elements in street-level imagery befor...
Abbas Salehitangrizi, S. Jabari, Yun Zhang· IEEE Journal of Selected Top...· 0 citations
This paper presents RICO-3D (Roadway Infrastructure in Context), a new large-scale Mobile Laser Scanning (MLS) dataset for semantic segmentation of French urban roadways, together with GA-Attention, a geometry-aware attention U-Net designed for this task. RICO-3D was acquired with a Leica Pegasus TRK300 mobile mapping...