Vision-Language Model–Driven Crack Detection for UAV-Based Inspection of High-Rise Infrastructure
Abstract
The structural integrity of aging high-rise concrete infrastructure requires frequent, high-precision inspections, yet traditional manual methods remain heavily reliant on labor-intensive, costly, and inherently dangerous techniques such as gondolas or rope access. This research aims to address these critical industry challenges by developing an innovative, autonomous cyber-physical framework that integrates Unmanned Aerial Vehicle (UAV) tactile scanning with advanced Vision-Language AI Models (VLM). The methodology employs an RTK-enabled quadcopter for high-fidelity data acquisition, processing 8,500 expert-annotated images through a multimodal transformer architecture that directly translates visual surface patterns into formal engineering narrative reports. The primary novelty of this work lies in the end-to-end automation of the diagnostic pipeline, which explicitly bridges the gap between raw pixel-level detection and professional semantic structural reasoning through the integration of the Crack Severity Index (CSI) and Serviceability Structural Performance (SSP) formulas. Results demonstrate that the proposed VLM framework achieves a superior detection accuracy with an mAP@0.5 of 0.962, consistently outperforming referenced convolutional neural network baselines. Furthermore, a comprehensive economic feasibility analysis reveals that this automated inspection system reduces total operational costs by 86.32% and decreases field inspection duration by 85.71% compared to conventional scaffolding-based techniques. This research contributes a robust, actionable diagnostic tool that significantly enhances safety standards and operational efficiency in modern urban asset management, effectively eliminating the subjectivity inherent in manual damage assessments while providing quantifiable data for long-term structural health monitoring.