A Comparative Investigation of YOLO26 and RF-DETR for Thin Crack Detection and Segmentation in UAV-Derived Airport Pavement Orthophotos
Abstract
Automated airport pavement inspection requires reliable instance segmentation models for detecting and quantifying thin cracks under real operating conditions. Building on a previously established UAV-AI workflow for airport pavement crack detection and quantification, and on an earlier investigation of sealed-crack class definition using YOLO11, the present study addresses a subsequent research question by comparing two recent model configurations based on fundamentally different computer vision paradigms. YOLO26 was selected for its deployment-oriented convolutional architecture, computational efficiency, and mechanisms aimed at improving small-target handling, whereas RF-DETR was selected for its transformer-based architecture and DINOv2-pretrained backbone; these provide fine-grained visual representations and exploit broader contextual information. The two configurations were assessed using the same dataset of 24,768 annotated images and compared in terms of computational demands, independent test-set performance, and field-based crack length reliability. Field validation was performed on two airport taxiways representing different surface conditions: taxiway Nibbio, mainly affected by active longitudinal and transverse cracks with limited interference from sealed cracks, and taxiway November, characterized by the coexistence of active and sealed cracking patterns. YOLO26 showed a lower computational demand, requiring approximately one hour and 20 compute units, compared with approximately six hours and 80 compute units for RF-DETR. RF-DETR achieved a higher mAP50 and recall on the test set and lower model error index values on both taxiways, indicating better crack length recovery. However, on taxiway November, it also showed higher hallucination index values, revealing greater sensitivity to visually ambiguous sealed cracks. These findings indicate that model selection should consider pavement surface conditions, computational constraints, and the operational consequences of missed cracks and false-positive detections. The specific contribution of the present study is therefore the extension of the previously established UAV-AI framework from workflow development and class definition analysis to the comparative evaluation of recent convolutional and transformer-based model configurations under real airport pavement conditions.