Aug 2026· International Bulletin of Applied Sciences and Technology· Vol 6, pp. 10-24· 0 citations
TL;DR
A Lightweight Transformer-Fourier Fusion Framework for Efficient Image Super-Resolution, designed to integrate the long-range dependency modeling capability of transformers with the frequency-domain representation advantages of Fourier-based feature processing.
Abstract
Image super-resolution (ISR) has emerged as a critical computer vision task aimed at reconstructing high-resolution visual information from low-resolution inputs. Although deep learning-based approaches have significantly improved reconstruction quality, many existing architectures suffer from high computational complexity, excessive parameter requirements, and limited efficiency in real-time deployment scenarios. This research presents a Lightweight Transformer-Fourier Fusion Framework for Efficient Image Super-Resolution, designed to integrate the long-range dependency modeling capability of transformers with the frequency-domain representation advantages of Fourier-based feature processing. The proposed framework is theoretically positioned around efficient feature extraction, adaptive attention learning, and frequency-aware reconstruction. Transformer-based modules enhance spatial relationship modeling, while Fourier convolution mechanisms improve the preservation of high-frequency image details with reduced computational overhead. The methodology combines lightweight residual feature refinement, transformer-driven contextual enhancement, and frequency-domain fusion to achieve an optimized balance between reconstruction accuracy and computational efficiency. The study analyzes the limitations of conventional convolutional, generative, and attention-based super-resolution approaches and establishes the importance of hybrid architectures for next-generation ISR systems. The proposed framework provides a scalable solution for applications requiring efficient image enhancement, including mobile imaging, medical visualization, remote sensing, and intelligent surveillance systems.
Image super-resolution (SR), which aims to reconstruct a high-resolution image from a low-resolution input, has progressed from convolutional neural networks (CNNs) to transformer-based architectures. Despite this progress, lightweight transformer SR remains challenging: local or window-based operations provide limited...
Single image super-resolution aims to reconstruct high-resolution images from low-resolution inputs. This paper proposes FreeTransformSR, a novel lightweight super-resolution network based on a channel-wise free low-rank learnable transform. The transform learns task-adaptive basis functions in a data-driven manner, en...
This paper proposes Adaptive ODE-ResNet (Adaptive ODE-ResNet), which reconstructs the residual block into a continuous ODE-driven process to realize flexible and accurate feature evolution and provides an accurate, efficient and scalable continuous-time modeling scheme for high-resolution image reconstruction.
Hai-Ying Zhang, Liangping Tu· Journal of King Saud Univers...· 0 citations
Convolutional neural networks (CNNs) have long played a central role in computer vision due to their fast computation speed and high image feature extraction performance. CNNs can effectively learn diverse visual features ranging from low-level to high-level representations, resulting in high computational efficiency....
This work proposes a lightweight dual-domain attention aggregation network (LDANet), aiming to achieve image super-resolution with both high efficiency and high quality, and proposes the pixel-embedding channel attention module, which achieves cross-channel global context awareness by jointly modeling pixel-level spati...
Wei Xue, Meng-Cheng Ma, Bing-Wen Hu et al.· ACM Transactions on Multimed...· 0 citations
A transformer-based channel attention block improves the discriminative capability of fused features in low-texture regions and enhances global consistency in a lightweight deep stitching framework that integrates multi-scale feature fusion with attention-enhanced matching.