2026· IEEE Transactions on Geoscience and Remote Sensing· Vol 64, pp. 3002422-3002422· 0 citations· 67 references
Abstract
The main bottleneck of current object detectors does not lie in weak feature representation. It lies in the lack of structured reasoning. Most detectors follow a recognition-style pipeline. They map features directly to detection results. This design ignores structural relations inside visual information. As a result, model stability and robustness degrade in complex scenes. This article proposes edge-frequency chain-of-thought (EF-CoT), a structured reasoning paradigm for object detection. EF-CoT reformulates detection as a stepwise reasoning process. It jointly models domain-specific tokens, spatial priors, and progressive reasoning. The model analyzes visual evidence before making predictions. Based on EF-CoT, this article designs the triple-branch guided frequency alignment and edge-aware network (TGFAENet). TGFAENet builds edge-frequency-aware structured tokens. Token-level reasoning then models the interaction among global structure, local details, and boundary cues. The Gaussian attention spatial pyramid pooling fusion (GASPPF) module is further introduced. GASPPF uses dual-branch spatial pyramid pooling and Gaussian attention (GA) to model continuous spatial priors. This module guides the detector toward potential target regions. By combining structured token modeling, spatial awareness, and chain-style reasoning, the proposed method builds a unified perception–analysis–decision detection framework. Experiments on complex remote sensing scenes show clear performance gains. The gains are more visible on small-object detection. These results support the shift from direct recognition to diagnostic reasoning in object detection.
Rank-Consistent Set Reasoning is presented, a supervised dense-prediction framework that models a group as an unordered set rather than as a sequence of images or a semantic label, and introduces a group permutation objective and hard-distractor augmentation so that the model learns the properties of a set-level target...
Lightweight oriented object detection in high-resolution remote sensing imagery is challenging, since detectors must handle substantial variations in object scale and complex object geometries under strict computational constraints. Prevailing lightweight methods often focus on backbone compression, leaving the neck an...
Remote sensing change detection still faces several challenges in complex scenes, including boundary ambiguity, cross-temporal appearance interference, and large object-scale variations. To address these issues, this article proposes LTS-ProtoNet, a local–temporal–semantic prototype network driven by a learnable change...
Real-time object detection is a cornerstone of intelligent transportation systems, where the YOLO (You Only Look Once) family has long defined the accuracy–speed trade-off. The emergence of Mamba—a selective state-space model that captures long-range dependencies in linear time—has triggered a wave of research embedd...
T. Vo, René Jaros, Ly Duc Minh et al.· Artificial Intelligence Revi...· 0 citations
A hyper look-ahead network is proposed, which incorporates a look-ahead structure (LS), conspicuous feature supplement attention (CFSA), and multiscale feature information process module (MFIPM) in the neck, which outperforms many state-of-the-art object detection methods.