Skip to content
Open access

Enhancing Image Inpainting Using Hybrid Architecture Combining CNN-Vision Transformer and GAN

Sep 2026 · Brain: Broad Research in Artificial Intelligence and Neuroscience · 0 citations

TL;DR

Both numerical results and visual evaluations confirm that the proposed framework provides a balanced solution for preserving structural details while maintaining perceptual realism in image inpainting applications.

Abstract

Image inpainting focuses on restoring missing parts of an image in a way that preserves both visual continuity and semantic consistency with the surrounding regions. In this study, a hybrid reconstruction model integrating convolutional neural networks, a Vision Transformer (ViT), and adversarial learning is presented to improve image completion quality. The convolutional layers are responsible for extracting local structural features, whereas the ViT component captures broader contextual relationships and long-range dependencies within the image. To enhance reconstruction performance, a combined loss structure including reconstruction, perceptual, style, edge, and adversarial losses was employed. The proposed approach was tested under various masking conditions through both quantitative and qualitative analyses. The experimental findings indicate that the model achieves strong reconstruction performance, reaching PSNR, SSIM, and LPIPS values of 36.22, 0.9779, and 0.0260, respectively. Additional ablation experiments demonstrate the importance of each architectural component. In particular, excluding the ViT module led to a noticeable decrease in PSNR performance, highlighting the significance of global contextual modeling. Likewise, removing the perceptual loss negatively affected perceptual similarity by increasing the LPIPS score. Visual evaluations indicated that the adversarial learning component contributed to generating more natural, visually coherent, and perceptually realistic reconstructions. Overall, both numerical results and visual evaluations confirm that the proposed framework provides a balanced solution for preserving structural details while maintaining perceptual realism in image inpainting applications.

Read PDF

Similar papers

Lightweight Diffusion‑GAN for Low‑Resource Image Inpainting

The method employs a lightweight UNet architecture based on depthwise separable convolutions, substantially reducing the parameter count while decreasing inference steps from 1000 to 20 through a fast sampling strategy, exhibiting excellent quality-efficiency trade-offs.

Yue Yu · 0 citations
Open access Aug 2026

PixelBoost 8 – Pixel Quality with 8X Highlights Boosting Enhancement

This work proposes a new GAN-based face hallucination method primarily based on the Enhanced Super-Resolution Generative Adversarial Network (ESRGAN), and presents a personalised adaptation of ESRGAN that employs the VGG16 architecture with a compact pre-trained version.

Sheetal S. Patil, A. Pawar, Nilofar Mulla et al. · 0 citations
Conference Open access Sep 2026

COMPARATIVE ANALYSIS OF DEEP LEARNING ARCHITECTURES FOR MONOCHROMATIC IMAGE COLORIZATION

Colorizing monochromatic images is crucial for enhancing the informativeness of visual data in modern information technology and computer vision systems. However, automatic colorization poses an inherently ill-posed mathematical challenge due to the multimodal nature of color distributions, where multiple valid color m...

Ihor Panasenko, K. Smelyakov, A. Chupryna et al. · 0 citations
Open access Aug 2026

Deep Edge-Aware Post-Processing for JPEG Enhancement: CNN-Based Artifact Reduction and Image Quality Restoration

A CNN-based edge-aware artifact reduction framework (CNN-AR) is proposed that integrates an enhanced deep super-resolution (EDSR) backbone with a holistically nested edge detection (HED) guided loss, enabling superior artifact suppression while preserving fine structural details.

Nupur, Nishant Kumar, Sajal Suhane et al. · 0 citations
2026

Person-Prioritized Restoration for High-Compression 360° Video

High-Compression videos suffer from severe distortions, among which degradation in person regions has the greatest impact on viewers’ immersive experience. Existing quality enhancement techniques usually focus on overall image denoising or super-resolution, often overlooking the crucial recovery of fine structures in t...

Lin-Yun Liu, Li Yu, Jia-Xin Zeng et al. · 0 citations
Open access Aug 2026

A Novel Image Inpainting Model Based on Multi-scale Parallel Dense Connection Network

A novel image inpainting framework based on a Multi-Scale Parallel Dense Connection Network (MSPDCN) with holistically nested edge detection first employed to extract structural priors and estimate edge information of missing regions, which provides guidance for subsequent reconstruction and alleviates boundary blurrin...

Jie Wang, Li-Yuan Zhang, Yi-Bo Deng et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.