Adaptive Frequency-Aware Diffusion (AFID), an adaptive framework for high-quality text-guided image inpainting, is proposed and comprehensive comparisons against prevailing state-of-the-art inpainting approaches solidly confirm its superiority in visual fidelity, intact-area preservation and text-prompt consistency.
Abstract
Image inpainting is a fundamental task in computer vision and multimedia processing. With the rapid development of denoising diffusion models, text-guided image inpainting has gradually become a mainstream research direction, enabling flexible content creation and localized semantic editing. Although existing text-guided diffusion inpainting approaches have greatly improved the generation quality and prompt alignment, they still face a core challenge: effectively balancing the fidelity preservation of unmasked regions and the semantic consistency of generated content in masked regions, especially suppressing spectral discontinuities and boundary artifacts caused by unreasonable frequency control. To address these problems, we propose Adaptive Frequency-Aware Diffusion(AFID), an adaptive framework for high-quality text-guided image inpainting. First, we design an Adaptive Frequency Threshold Network (AFTN) to dynamically predict multi-band frequency cutoffs according to mask information, denoising timesteps and text conditions, replacing manually designed fixed rules. Second, we propose a two-stage smooth spectral blending strategy to alleviate spectral abruptness and enhance the natural coherence between masked and unmasked areas. Third, we introduce a boundary ring smoothing module to further eliminate stitching artifacts near mask edges without excessive blurring. Experimental results show that our method achieves a leading performance on standard public benchmarks. We conduct comprehensive comparisons against prevailing state-of-the-art inpainting approaches, which solidly confirm our superiority in visual fidelity, intact-area preservation and text-prompt consistency. Extensive ablation studies verify the necessity of all three core modules and further analyze the impact of frequency thresholds and blending configurations on final restoration outcomes.
A Structure-Guided Textual Mask Network is designed to predict geometry-aware editing regions by leveraging refined textual structural cues and human-centric priors, where a structural prior reweighting mechanism is introduced to improve localization accuracy.
Xin Chen· Poster Volume 0007 The 2026...· 0 citations
This work proposes a subject clarity outpainting framework that combines vision-language model (VLM)-guided semantic conditioning with multiscale wavelet supervision for subject-localized detail preservation and develops a subject-centric data curation pipeline that constructs subject-intersecting outpainting pairs fro...
Abhilash Neog, Taewan Kim, Yi Wu et al.· 0 citations
Diffusion models with prompt and reference image-guided editing have seen rapid progress, yet they remain too coarse for pixel-level control. One promising direction is to incorporate a soft mask that specifies spatially varying edit strengths but training such fine-grained control demands expensive pixel-wise annotati...
MagnifiQ is introduced, an image restoration framework that progressively upscales and restores images across resolutions, e.g., from 1024x1024 to 4096x4096, and leverages a pre-trained text-to-image diffusion model such as SDXL and adapts it for more scalable high-resolution inference by replacing its original self-at...
M. Reddy, Yashesh Savani, Antoine Mercier et al.· 0 citations
A novel image inpainting framework based on a Multi-Scale Parallel Dense Connection Network (MSPDCN) with holistically nested edge detection first employed to extract structural priors and estimate edge information of missing regions, which provides guidance for subsequent reconstruction and alleviates boundary blurrin...
Jie Wang, Li-Yuan Zhang, Yi-Bo Deng et al.· 電腦學刊· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.