A transformer-based channel attention block improves the discriminative capability of fused features in low-texture regions and enhances global consistency in a lightweight deep stitching framework that integrates multi-scale feature fusion with attention-enhanced matching.
Abstract
Image stitching aims to construct wide field-of-view scenes from multiple narrow-FoV images, yet existing deep learning-based approaches may introduce substantial computational and parameter overhead, limiting their applicability in efficiency-sensitive scenarios. To address this issue, we propose a lightweight deep stitching framework that integrates multi-scale feature fusion with attention-enhanced matching. Specifically, a transformer-based channel attention (TCA) block improves the discriminative capability of fused features in low-texture regions and enhances global consistency. A coordinate-aware correlation module (CACM) combines correlation-based matching with position-sensitive coordinate attention to support registration under parallax, while GhostNet serves as the shared backbone. On UDIS-D, the complete model achieves a 25.19 dB peak signal-to-noise ratio (PSNR) and a structural similarity index (SSIM) of 0.833 under the overlap-region protocol, with a reported full-system complexity of 19.28 giga multiply-accumulate operations (GMACs) and 52.86 M parameters. These results demonstrate a favorable accuracy–efficiency trade-off under the stated GMAC, runtime, and memory protocol and the potential of the proposed framework for resource-conscious image stitching applications.
Aiming at the problems of large parameters and high computational complexity in deep learning-based image super-resolution networks, this paper proposes a lightweight super-resolution network that fuses multi-attention mechanism and Blueprint Separable Convolution (BSConv). BSConv is introduced to improve performance w...
Yi-Yan Huang, Lin Guo· International Conference on...· 0 citations
High Dynamic Range (HDR) reconstruction from multi-exposure Low Dynamic Range (LDR) images requires recovering a wide luminance range while preserving details in bright and dark regions under motion and exposure misalignment. High reconstruction fidelity, however, often comes with increased computational complexity. Th...
Ian Oliveira Teixeira, Q. Leher, Josue Lopez-Cabrejos et al.· Pattern Analysis and Applica...· 0 citations
Image deblurring is a challenging task in computer vision because it is a difficult and spatially variant problem. This work presents a Transformer-based architecture that utilizes self-attention mechanisms to capture useful long-range dependencies in an image. Unlike traditional convolutional approaches, the proposed...
Sudharshan Banakar, K. Chandrashekhar· Engineering, Technology &...· 1 citation
A dense matching network based on a Transformer and multi-scale feature fusion, called Task-aware Multi-Scale Matching Network (TMSMNet) is proposed, which outperforms mainstream methods such as RAFT-Stereo on the D1-all metric of KITTI- 2015 and demonstrates good generalization and robustness.
Shi-Xiong Liu· ITM Web of Conferences· 0 citations
This work proposes SLRNet, Super Lightweight Residual Network, a high efficient-yet-effective end-to-end dehazing architecture that integrates a novel Adaptive Feature Unit that automatically adjusts channel-wise features through a lightweight gating mechanism, coupled with compact residual blocks to preserve critical...
Guanheng Qu· Poster Volume 0007 The 2026...· 0 citations
A Lightweight Transformer-Fourier Fusion Framework for Efficient Image Super-Resolution, designed to integrate the long-range dependency modeling capability of transformers with the frequency-domain representation advantages of Fourier-based feature processing.
Chinedu Okafor· International Bulletin of Ap...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.