Analysis of multichannel satellite images of urban areas using deep learning methods
This paper examines the problem of land cover classification in densely populated urban environments using ultra-high-resolution Earth observation images. The aim of the study is to develop and validate a neural network algorithm for a high-precision semantic segmentation of urban green spaces using multimodal data. An improved U-Net convolutional neural network architecture, modified for the use with a 7-channel input tensor (RGB, NIR, RedEdge, DSM, and NDVI), is proposed. The approach is based on the Early Fusion strategy, which combines spectral measurements with promising structural characteristics (digital surface model, DSM). Focal Loss is used to overcome the class imbalance. The proposed model reliably separated spectrally identical layers (grass and trees) and eliminated false positives on green anthropogenic objects. The final Mean IoU was 0.725. The recall for detecting forested areas reached 0.93, and for the complex minority class “Shrubs” it reached 0.77. The experiment on an independent test site confirmed the model’s high generalizability (F1-score 0.93). The integration of seven data channels and the use of a modified U-Net are fully justified for the tasks of accurate calculating forest areas and environmental monitoring in the Smart City concept.