Skip to content
Open access

Dual-Path Collaborative Transformer: Hybrid Attention and Stepwise Dilated Convolution for Remote Sensing Image Super-Resolution

Sep 2026 · Italian National Conference on Sensors · Vol 26, pp. 5600 · 0 citations · 54 references
Medicine

TL;DR

A novel dual-path collaborative architecture, named DCTNet, which combines Transformer-based global modeling with convolution-driven local feature extraction and achieves competitive reconstruction performance across different scale factors, with statistically significant improvements observed in specific settings.

Abstract

Remote Sensing Image Super-Resolution (RSISR) is a core task in geospatial image analysis. Convolutional neural networks (CNNs) have achieved significant breakthroughs in RSISR tasks by extracting local features. However, CNN-based methods struggle to capture long-range dependencies, thereby limiting SR performance. Recently, Transformer-based methods have demonstrated remarkable performance in capturing global information. Nevertheless, they remain inadequate for exploring high-frequency details and local features. To overcome these limitations, this work introduces a novel dual-path collaborative architecture, named DCTNet, which combines Transformer-based global modeling with convolution-driven local feature extraction. DCTNet is a hybrid network composed of a CNN-Transformer Residual Hybrid Group (CTHG). This group consists of two core components: the Dual-domain Fusion Window Attention Block (DFWAB) and the Stepwise Dilated Convolution (SDC). Specifically, the DFWAB incorporates channel and frequency attention mechanisms following the standard Transformer block to recover high-frequency details. Furthermore, by integrating stepwise dilated convolutions into the conventional Transformer architecture, the CTHG effectively captures both multi-scale local and global features. Additionally, we employ dense connections among the DFWAB modules to facilitate feature reuse across layers. Experimental results on the AID and UCMerced datasets demonstrate that DCTNet achieves competitive reconstruction performance across different scale factors, with statistically significant improvements observed in specific settings.

Read PDF

Similar papers

Open access 2026

EDG-Net: A Lightweight Frequency-Aware CNN-Transformer Hybrid Network for Efficient Remote Sensing Change Detection

Accurate change detection (CD) in high-resolution remote sensing imagery is often affected by “pseudochange” interference (e.g., seasonal and illumination variations) and by the computational constraints of edge devices. To address these challenges, we propose the efficient difference-gated network (EDG-Net), a lightwe...

Qing-Xiang Meng, Jin-Ning Zhao, Wen-Jie Yue et al. · 0 citations
Open access Sep 2026

Lightweight Dual-Domain Attention Aggregation Network for Remote Sensing Image Super-Resolution

Significant progress has been made in remote sensing image super-resolution based on deep neural networks. However, existing methods typically suffer from parameter redundancy and high computational costs, making them difficult to deploy on resource-constrained edge devices. Moreover, the image reconstruction process o...

Wei Xue, Meng-Cheng Ma, Bing-Wen Hu et al. · 0 citations
Sep 2026

ALSRFormer: An adaptive transformer with dynamic window attention and multi-scale deformable feed-forward network for remote sensing image segmentation.

High-resolution remote sensing images present considerable challenges for semantic segmentation due to their complex object structures and extensive spatial distribution. Effective segmentation requires capturing fine-grained local details while simultaneously modeling long-range dependencies. Convolutional Neural Netw...

Zi-Qi Jia, Jia-Wei Zhang, Dong-En Guo et al. · 0 citations
Open access Aug 2026

DAAF: Dual-stream adaptive attention fusion with distribution alignment for remote sensing object detection

Remote sensing object detection faces three challenges: extreme scale variation, arbitrary rotational orientations, and complex intermingled backgrounds. Although fusion of Convolutional Neural Networks (CNNs) and Transformers combines spatial precision with global modeling, it faces two limitations: (i) feature distri...

Mo Zhou, Yue Zhou, Kai Song · 0 citations
Sep 2026

Position-Aware Dual-Domain Hybrid Transformer for Single Hyperspectral Image Super-Resolution

Hyperspectral image super-resolution (HSI-SR) is a technique that increases spatial resolution while preserving spectral accuracy. This capability is important for remote sensing interpretation and for many downstream applications. Most existing deep learning solutions are predominantly built upon convolutional neural...

Hai-Jun Wang, Hao-Yu Hu, Ya-Lin Nie et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.