Skip to content
Open access

SVI2LoD3: Agent-Driven Reconstruction of LoD3 Façade Openings in Semantic 3D City Models from Volunteered Street View Imagery using Large Language and Visual Models

Aug 2026 · ISPRS Annals of the Photogrammetry, Remote Sensing and Spatial Information Sciences · 1 citation · 9 references
Computer Science

TL;DR

An end-to-end, agent-driven pipeline for the LoD3 reconstruction of façade openings in 3D city models, producing directly usable CityGML-conform outputs and a novel evaluation metric for façade reconstruction, termed Facade Feature Distance.

Abstract

This paper presents an end-to-end, agent-driven pipeline for the LoD3 reconstruction of façade openings in 3D city models, producing directly usable CityGML-conform outputs. In contrast to existing approaches that rely on supervised semantic segmentation and therefore require large amounts of manually annotated training data, the proposed method employs a zero-shot segmentation strategy. This substantially reduces the annotation effort while still achieving strong performance in our benchmark on the eTRIMS dataset. A further key contribution is the enforcement of correct partonomic hierarchies, thereby producing CityGMLconform LoD3 building models. Beyond the reconstruction pipeline itself, this work also introduces a novel evaluation metric for façade reconstruction, termed Facade Feature Distance (FFD). Unlike conventional metrics such as mIoU or FRDS, which assess similarity primarily through pixel-wise overlap, FFD measures distance in a high-level feature space derived from a vision transformer. In doing so, it captures both semantic correctness and architectural layout, providing a more suitable assessment of façade reconstruction quality. The proposed pipeline and evaluation strategy together offer a practical and scalable contribution toward the automated generation and analysis of semantically enriched 3D city models. The developed code is published at: https://github.com/hcu-cml/citydb-SVI2LoD3-ai.

Read PDF

Similar papers

Open access Sep 2026

Omni2LOD3: Extracting Façade Geometry and Semantic Features from Omnidirectional Imagery and Drone Photogrammetry

CityGML Level of Detail 3 (LOD3) building model generation requires accurate façade representation. Existing workflows often rely on LiDAR, dense multi-view imagery, or pre-existing lower-LOD models, limiting their applicability in data-scarce environments. Additionally, most methods are developed on buildings with sim...

Demi Julianna L. Gentiles, Khalil C. Torneros, J. Casisirano et al. · 0 citations
Review Open access Sep 2026

End-to-End Pothole Detection and Semantic 3D City Model Enrichment Using LoRA-Adapted Foundation Models and Open Street-Level Imagery

Regular road inspection is essential for maintaining pavement service life and ensuring traffic safety. Conventional approaches require specialized survey vehicles and sensor payloads, while damage assessment relies predominantly on supervised computer vision models trained on large annotated datasets. Recent advances...

B. Ouzougarh, Jannik Matijevic, Huynh Duc An Son Nguyen et al. · 0 citations
#artificial intelligence Review Aug 2026

Polis: 3D Self-Supervision at City Scale

The results show that distributionally-regularized joint embedding architectures can be successful on challenging city-scale 3D scenes, and that transfer improves when self-supervision is designed for the capture geometry and spatial context of this domain while also revealing the limits of this specialization.

A. Rusnak, S. Kovalenko, Jingru Wang et al. · 0 citations
Preprint Oct 2026

Lang3DSeg: Annotation-Free Open-Vocabulary 3D Segmentation with Point Transformers

Accurate 3D semantic perception is critical for safe autonomous navigation. However, supervised LiDAR segmentation remains tied to closed taxonomies and to the cost of point-wise manual annotation. Open-vocabulary methods avoid that cost by projecting the output of 2D vision-language models onto LiDAR and distilling it...

Cigdem Kokenoz, Amir Salarpour, Alkim Domeke et al. · 0 citations
Open access 2026

OccluFree: Occlusion-Aware Restoration for Amodal Building Windows and Doors Segmentation

Urban applications, such as facade reconstruction, digital twins, building analysis, and semantic 3-D city modeling benefit from reliable and geometrically consistent representations of window and door (WD) elements. Many image-based facade reconstruction pipelines first detect WD elements in street-level imagery befor...

Abbas Salehitangrizi, S. Jabari, Yun Zhang · 0 citations
Open access Aug 2026

RICO-3D: A Benchmark and Baseline Method for Semantic Segmentation of Urban Roadways

This paper presents RICO-3D (Roadway Infrastructure in Context), a new large-scale Mobile Laser Scanning (MLS) dataset for semantic segmentation of French urban roadways, together with GA-Attention, a geometry-aware attention U-Net designed for this task. RICO-3D was acquired with a Leica Pegasus TRK300 mobile mapping...

Wided Hammedi, Olivier Hotel, Franck Roudet et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.