Skip to content

From Recognition to Reasoning: Edge-Frequency Chain-of-Thought for Remote Sensing Object Detection

2026 · IEEE Transactions on Geoscience and Remote Sensing · Vol 64, pp. 3002422-3002422 · 0 citations · 67 references

Abstract

The main bottleneck of current object detectors does not lie in weak feature representation. It lies in the lack of structured reasoning. Most detectors follow a recognition-style pipeline. They map features directly to detection results. This design ignores structural relations inside visual information. As a result, model stability and robustness degrade in complex scenes. This article proposes edge-frequency chain-of-thought (EF-CoT), a structured reasoning paradigm for object detection. EF-CoT reformulates detection as a stepwise reasoning process. It jointly models domain-specific tokens, spatial priors, and progressive reasoning. The model analyzes visual evidence before making predictions. Based on EF-CoT, this article designs the triple-branch guided frequency alignment and edge-aware network (TGFAENet). TGFAENet builds edge-frequency-aware structured tokens. Token-level reasoning then models the interaction among global structure, local details, and boundary cues. The Gaussian attention spatial pyramid pooling fusion (GASPPF) module is further introduced. GASPPF uses dual-branch spatial pyramid pooling and Gaussian attention (GA) to model continuous spatial priors. This module guides the detector toward potential target regions. By combining structured token modeling, spatial awareness, and chain-style reasoning, the proposed method builds a unified perception–analysis–decision detection framework. Experiments on complex remote sensing scenes show clear performance gains. The gains are more visible on small-object detection. These results support the shift from direct recognition to diagnostic reasoning in object detection.

View source

Similar papers

#small language model Preprint Sep 2026

Rank-Consistent Set Reasoning for Co-Salient Object Detection

Rank-Consistent Set Reasoning is presented, a supervised dense-prediction framework that models a group as an unordered set rather than as a sequence of images or a semantic label, and introduces a group permutation objective and hard-distractor augmentation so that the model learns the properties of a set-level target...

Yuan-Xin Xiang, Matteo Rossi, Ying-Zhou Chen · 0 citations
2026

IBODet: Information-Bottleneck-Inspired Lightweight Oriented Detection in Remote Sensing Images

Lightweight oriented object detection in high-resolution remote sensing imagery is challenging, since detectors must handle substantial variations in object scale and complex object geometries under strict computational constraints. Prevailing lightweight methods often focus on backbone compression, leaving the neck an...

Ling-Xiang Hao, Ling-Fei Hao, Chubo Deng · 0 citations
Open access 2026

LTS-ProtoNet: Local–Temporal–Semantic Prototype Network for Remote Sensing Change Detection

Remote sensing change detection still faces several challenges in complex scenes, including boundary ambiguity, cross-temporal appearance interference, and large object-scale variations. To address these issues, this article proposes LTS-ProtoNet, a local–temporal–semantic prototype network driven by a learnable change...

Hui Zhang, Zhao-Long Gao, Yan-Yan Mao · 0 citations
Review Open access Aug 2026

Selective state-space modeling for real-time object detection in intelligent transportation: a systematic literature review

Real-time object detection is a cornerstone of intelligent transportation systems, where the YOLO (You Only Look Once) family has long defined the accuracy–speed trade-off. The emergence of Mamba—a selective state-space model that captures long-range dependencies in linear time—has triggered a wave of research embedd...

T. Vo, René Jaros, Ly Duc Minh et al. · 0 citations
Open access 2026

Hyper Look-Ahead Network for Remote Sensing Object Detection

A hyper look-ahead network is proposed, which incorporates a look-ahead structure (LS), conspicuous feature supplement attention (CFSA), and multiscale feature information process module (MFIPM) in the neck, which outperforms many state-of-the-art object detection methods.

Linfeng Jiang, Ya-Hao Li, Ting Bai · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.