Aug 2026· Expert Syst. J. Knowl. Eng.· Vol 43· 0 citations· 98 references
Computer Science
TL;DR
This survey systematically categorizes the 30‐year evolution of image inpainting into three distinct technological generations: traditional prior‐driven synthesis, deep learning data‐driven reconstruction and modern foundation model‐driven generation, providing a definitive reference for future theoretical and engineering advancements.
Abstract
As a fundamental task in restoring continuous visual signals, image inpainting plays a critical role in autonomous driving perception, medical imaging, video editing and digital heritage preservation. Driven by deep learning and large‐scale generative models, the field has transitioned from low‐level texture synthesis to high‐level semantic generation, yielding major breakthroughs in structural fidelity and visual realism. Centring on the generative paradigm as the architectural trajectory, this survey systematically categorizes the 30‐year evolution of image inpainting into three distinct technological generations: traditional prior‐driven synthesis, deep learning data‐driven reconstruction and modern foundation model‐driven generation. Despite this progress, highly competitive methods still struggle with large‐scale missing regions, global consistency in complex scenes, fine‐grained micro‐details and alignment with human visual perception. To address these gaps, we critically evaluate the technical paradigms and main bottlenecks within each of these evolutionary stages. We categorize and compare mainstream breakthroughs across high‐resolution restoration, text‐guided synthesis and complex scene generation. Furthermore, we compile standard benchmarks, evaluation metrics and quantitative performance comparisons of representative algorithms. Finally, we dissect open challenges—focusing on cross‐scene generalization and evaluation metric alignment—and outline future trajectories, particularly the integration of inpainting with text‐guided foundation models, providing a definitive reference for future theoretical and engineering advancements.
Adaptive Frequency-Aware Diffusion (AFID), an adaptive framework for high-quality text-guided image inpainting, is proposed and comprehensive comparisons against prevailing state-of-the-art inpainting approaches solidly confirm its superiority in visual fidelity, intact-area preservation and text-prompt consistency.
Both numerical results and visual evaluations confirm that the proposed framework provides a balanced solution for preserving structural details while maintaining perceptual realism in image inpainting applications.
Simge Coşkun, A. Işık· Brain: Broad Research in Art...· 0 citations
The method employs a lightweight UNet architecture based on depthwise separable convolutions, substantially reducing the parameter count while decreasing inference steps from 1000 to 20 through a fast sampling strategy, exhibiting excellent quality-efficiency trade-offs.
High-Compression videos suffer from severe distortions, among which degradation in person regions has the greatest impact on viewers’ immersive experience. Existing quality enhancement techniques usually focus on overall image denoising or super-resolution, often overlooking the crucial recovery of fine structures in t...
Lin-Yun Liu, Li Yu, Jia-Xin Zeng et al.· IEEE Signal Processing Lette...· 0 citations
An adaptive multi-scale decoding framework that effectively balances global context with fine-grained detail is proposed that exhibits superior robustness and generalization across diverse domains, effectively alleviating limitations of existing fusion-based approaches.