Efficient Hybrid Transformer-CNN Network for Real-World Image Denoising
Abstract
Convolutional neural networks (CNNs) have long played a central role in computer vision due to their fast computation speed and high image feature extraction performance. CNNs can effectively learn diverse visual features ranging from low-level to high-level representations, resulting in high computational efficiency. However, CNNs suffer from the loss of global contextual information, leading to loss of image detail and difficulty maintaining structural consistency in reconstructed images. Recently, Transformer-based models have been introduced in computer vision research. Transformers can model relationships between all locations within an input sequence, demonstrating their exceptional ability to capture long-range dependencies. However, as image resolution increases, memory usage and computational cost in image processing tasks increase drastically. To address the limitations of both CNN and Transformer, this paper proposes an efficient hybrid model that combines CNN and Transformer modules. The proposed model integrates Transformer and CNN modules to simultaneously leverage global contextual feature extraction by the Transformer and local fine-grained feature extraction by the CNN. Experimental results demonstrate that the proposed model achieves 39.22 dB PSNR and 0.946 SSIM on the SIDD dataset, and 39.41 dB PSNR and 0.949 SSIM on the DnD dataset, delivering competitive performance while significantly reducing computational cost.