GECNet: A Vectorized Building Footprints Extraction Network Based on Geometric Perception and Vertex Guidance
Abstract
Extracting building footprints from aerial or satellite imagery remains a significant challenge, particularly in maintaining the geometric regularity of man-made structures. While polygon-based methods offer vectorized representations superior to pixel-based approaches, they often struggle with corner ambiguity and fail to preserve structural constraints like parallelism and orthogonality. To address these limitations, we propose a geometry-aware framework for building footprint extraction that enforces geometric consistency through three coupled components: 1) a multitask network that synergistically learns semantics and geometry by jointly optimizing instance segmentation, vertex prediction, and boundary segments; 2) a geometry-aware mask generation module that refines boundaries by aligning semantic features with geometric cues; and 3) a vertex-guided polygon extraction module that reconstructs footprints while explicitly incorporating geometric constraints to rectify irregular shapes. Evaluated on the AICrowd, OpenCity, and Inria datasets, our method not only achieves notable improvements in standard metrics (AP, intersection over union (IoU), and PoLiS) but also produces vector outcomes with significantly higher geometric regularity compared to state-of-the-art methods.