Jul 2026· ISPRS Annals of the Photogrammetry, Remote Sensing and Spatial Information Sciences· Vol XI-2-2026, pp. 1-8· 0 citations· 17 references
TL;DR
Overall, state-of-the-art deep learning-based matchers still struggle with large rotations, scale differences, and semantic differences, and strongly benefit from prior image orientation knowledge and lack sub-pixel precision.
Abstract
Abstract. Accurate image matching is essential for the precise orientation of airborne imagery, yet modern feature matchers are rarely evaluated on real aerial data with great temporal, seasonal, and radiometric changes. For this study, we introduce the AerialRefMatch dataset, which comprises 51 challenging aerial images and corresponding true-ortho reference data. We benchmark classical and deep learning–based matching algorithms on AerialRefMatch, considering two scenarios: matching original images and matching approx-orthorectified images generated using GNSS/IMU orientations. For each method, image-based ground control points are derived and used for single-image pose estimation; accuracy is assessed via independent checkpoints. Results show that directly matching on original images is very difficult: fewer than 14% of images can be oriented with pixel-level accuracy. When approxorthorectification is used, performance improves substantially. JamMa, SIFT, and SuperPoint+LightGlue achieve pixel-level accuracy for up to 30% of images, with JamMa being most robust on difficult cases and SIFT-based variants being more precise on the easier ones. Deep detector-free models such as ELoFTR and RoMa are less accurate but more robust to the original images than other models. Overall, state-of-the-art deep learning-based matchers still struggle with large rotations, scale differences, and semantic differences, and strongly benefit from prior image orientation knowledge and lack sub-pixel precision. The AerialRefMatch dataset can be downloaded here: https://www.dlr.de/en/eoc/aerial-ref-match
Abstract. Archival images or videos collected for various purposes may be a valuable data source for photogrammetric processing. Many such examples are reported in the literature. However, the examples of the VHS aerial video photogrammetric processing are not mentioned. This investigation aims to evaluate whether such a type of data collected in a challenging scenario, i.e., an ad-hoc helicopter flight over a flooded area, can be processed with the SfM software. The analysis is focused on image matching and bundle block adjustment. The test data consists of 6 image sequences extracted from VHS videos collected with two different cameras: RGB and near infrared (NIR) from two different heights along the Odra river near Wroclaw during the 1997 Central European Flood. In addition to the images, a low accuracy GPS trajectory logs were used to extract approximated image positions to support tie-point extraction. Archival orthophotomap and digital surface model were used to create ground control points (GCPs) and check points. Each image sequence was processed separately in Agisoft Metashape Professional software. Results showed that the matching of VHS images is possible. Although the trajectory data was of very low accuracy, it was essential to perform tie-point extraction. GCPs extracted from other geospatial data allowed for image georeferencing and creation of orthomosaics, though the geometrical accuracy of the created products is rather low. Nevertheless, executed experiments proved the usefulness of such images, especially NIR, in the flooded area mapping, in which the photogrammetric processing can be executed with the commercial SfM software.
Grzegorz Jóźków, Maurycy Hechmann· The International Archives o...· 0 citations
Abstract. In the initial response to wildfires, securing rapid and accurate geographic information is essential. However, helicopter imagery acquired on-site often lacks precise sensor metadata, such as camera pose and internal parameters, making the application of georeferencing difficult. In particular, obliquely captured wildfire imagery presents additional registration challenges due to severe viewpoint changes, scale variations, and low-texture environments. This study proposes an automated georeferencing pipeline capable of operating under these constraints. The proposed method consists of five stages: preprocessing, image retrieval, feature extraction and matching, Exterior Orientation Parameters (EOP) estimation, and orthomosaic generation. An initial Area of Interest (AOI) is defined using inaccurate initial position data, and the Region of Interest (ROI) within the reference map is obtained through a ResNet50-based image retrieval approach. Subsequently, virtual Ground Control Points (GCPs) are generated through deep learning-based feature matching. Elevation data is then assigned using a Digital Elevation Model (DEM), and EOP are estimated via Perspective-n-Point (PnP) and RANSAC algorithms. Intermediate frames are initialized via interpolation and refined through bundle adjustment to produce the final orthomosaic. Experimental results demonstrated that utilizing SuperGlue and LightGlue complementarily increased the number of successfully georeferenced intervals from 5 to 9. Furthermore, a minimum RMSE of 28.30 m was achieved in the most accurate interval. This method proves that by automating the feature-based georeferencing process, practical geographic information can be rapidly provided for initial disaster response, even in sensor-limited environments.
Seongyun Kim, Jeonghyo Oh, J. Cheon et al.· The International Archives o...· 0 citations
The upgraded version of the geometric correction module of the STORM processing chain can automatically orthorectify images from the NEMO-HD small satellite, which, like other small satellites, in principle has a lower signal-to-noise ratio (SNR) and higher radiometric variability.
Aleš Marsetič, P. Pehani, Nina Krašovec· The International Archives o...· 0 citations
Abstract. Structure-from-Motion (SfM) pipelines rely heavily on the detection and matching of repeatable keypoints across images, yet the performance of modern learned feature extractors in challenging environments remains insufficiently understood. This paper evaluates classical and deep keypoint detectors for SfM reconstruction using winter Arctic UAV imagery, a domain characterized by low texture, repetitive patterns, and limited man-made structure. We compare three feature pipelines within a shared PyCOLMAP-based framework: SIFT with nearest-neighbor matching (SIFT+NN), SuperPoint, and DISK, along with a hybrid approach combining SuperPoint and DISK correspondences. Quantitative evaluation is conducted using standard SfM metrics, including number of observations, track length, observations per image, and reprojection error, complemented by qualitative analysis of keypoint distributions and reconstruction interpretability. Results show that SIFT+NN consistently achieves the most complete and stable reconstructions, producing the highest number of matched observations and lowest reprojection error across aggregate experiments. However, on more challenging subsets lacking clear structural features, learned methods demonstrate improved robustness, successfully reconstructing multiple views where SIFT fails. SuperPoint provides broader spatial coverage, while DISK produces denser clusters in high-confidence regions, highlighting complementary behaviors between learned approaches. Overall, the findings indicate that classical methods remain strong baselines for Arctic UAV photogrammetry under standard SfM pipelines, while learned detectors offer advantages in difficult conditions. The observed performance gap is attributed to domain mismatch and backend optimization for handcrafted features. These results suggest that domain-specific training and improved spatial feature distribution are promising directions for advancing learned keypoint methods in Arctic reconstruction tasks.
Nicholas Sansoterra, M. G. Lenzano, William J. Shuart et al.· The International Archives o...· 0 citations
Abstract. This paper focuses on cross-modal image matching between Synthetic Aperture Radar (SAR) and optical imagery, a longstanding challenge due to fundamental differences in imaging geometry and radiometry. Beyond applicational needs in satellite data fusion and downstream mapping, this study is motivated by the rapid advances in the field of Computer Vision. Thus, this work evaluates classical and modern learning-based feature matching methods on the renowned SpaceNet9 dataset using a unified evaluation framework. The results show that classical methods such as SIFT fail to produce reliable correspondences, while learning-based approaches, particularly MINIMA, significantly improve performance without additional retraining. However, matching accuracy is strongly influenced by scene structure and SAR-specific geometric effects, therefore robust SAR-optical correspondence remains an open challenge.
Constantin Günzel, M. Schmitt· The International Archives o...· 0 citations
Abstract. The radiometric adjustment of aerial imagery is a process of very high importance considering the influences this step can generate not only on the look of the image data (white balance), but even more importantly on derived information like indices (NDVI). In comparison to the Aerial Triangulation where it is relatively straight forward to set up thresholds that need to be met in order to achieve a high-quality result, the world of radiometric adjustment is dramatically different. There is no single standard or guideline that dictates what a high-quality radiometric result will look like. Apart from these challenges there is also a rather big gap between the rich and longstanding academic work done in the field of radiometry and actual application in real-life projects. The biggest discrepancy is the usage of single images, especially when dealing with absolute radiometry approaches versus multiple thousand color balanced images in a single block in an actual production environment. In this paper, we present the advantages of utilizing reflectance measurements as a method to stabilize radiometric adjustments, as well as utilizing them as anchor to create indices like the NDVI that correspond to the value range given by literature.
S. Scholz, M. Muick· The International Archives o...· 0 citations