Progressive Diffusion With Sparse Point Hypotheses for Moving Vehicle Detection From Satellite Videos
Abstract
Detecting moving vehicles in satellite videos remains challenging due to the extremely small target size and the severe foreground–background imbalance. Most existing approaches extend convolution-based dense detectors to the temporal domain and rely on fixed grid-aligned prediction, which becomes inefficient and unstable under highly sparse foreground conditions. In this work, we propose DiffMOD, a framework that reformulates moving vehicle detection as a diffusion-based denoising process over sparse point hypotheses, where candidate locations are progressively refined from noise toward target centers. To adapt diffusion modeling to the high-sparsity remote sensing scenario, a density-guided initialization mechanism is introduced to guide hypothesis sampling toward informative regions using lightweight density estimation. A spatial-relation self-attention (SRSA) module is further designed to enhance interactions among sparse hypotheses and aggregate contextual spatio-temporal cues for recovering weak targets in cluttered scenes. In addition, a coverage-constrained assignment (CCA) strategy combined with a progressive training scheme is developed to stabilize optimization and gradually tighten supervision during training. Extensive experiments on the RsCar and SDM-Car benchmarks demonstrate the effectiveness of the proposed method. DiffMOD achieves state-of-the-art performance and improves the previous best method by 6.3% in $F1$ -score and 13.8% in recall on the challenging SDM-Car dataset.