Skip to content
Conference

FE-GAN: feature enhancement and fine-grained interaction for text-to-image synthesis

Jul 2026 · International Conference on Generative Artificial Intelligence and Image Processing · Vol 14292, pp. 142920D - 142920D-11 · 0 citations · 15 references
Engineering

TL;DR

A GAN-based method for generating images from text by introducing multi-kernel convolution into the generator and replacing standard convolutions with multi-layer nested convolutions combined with dynamic gating weighting is proposed, thereby improving the fine-grained quality of generated images.

Abstract

In the task of text-to-image synthesis, it is challenging to generate semantically consistent and high-fidelity images from given text descriptions. To address the issues of semantic inconsistency between text and images and the lack of realism in generated image details, this paper proposes a GAN-based method for generating images from text. Specifically, by introducing multi-kernel convolution into the generator and replacing standard convolutions with multi-layer nested convolutions combined with dynamic gating weighting, the model is encouraged to focus on crucial detailed features, thereby improving the fine-grained quality of generated images. Then, we design a feature enhancement block that models multi-dimensional feature interactions via channel, spatial height, and spatial width branches, which effectively enhances feature representation and produces more photo-realistic images. Furthermore, we propose a fine-grained interaction mechanism that utilizes a low-rank correlation matrix to capture the dependencies between global and local information at different granularity levels, achieving fine-grained channel enhancement and significantly improving text-image alignment accuracy. Moreover, the proposed model maintains the advantages of efficient generation and a smooth latent space from the GAN paradigm. Finally, comprehensive experiments are carried out on the CUB and COCO benchmarks. Quantitative results illustrate that our approach achieves 11.48 FID and 5.23 IS on the CUB dataset, and 17.46 FID and 36.03 IS on the COCO dataset, respectively.

View source

Similar papers

Preprint Sep 2026

BTC3D: Blended Tile Conditioning for Detail-Enhancing Image-to-3D Generation

Recent diffusion-based pipelines have achieved promising progress in image-to-3D synthesis. However, generating high-fidelity details remains challenging, especially when the input image contains rich details. Existing approaches often rely on globally encoded conditioning features, which compress spatial information a...

Jun-Yu Li, Qiu-Yu Chen, Peng-Cheng Wang et al. · 0 citations
Preprint Aug 2026

PixelControl: Fine-Grained Condition Fidelity in Text-to-Image Diffusion

Experiments show that PixelControl improves structural fidelity and visual quality over existing controllable generation methods, with especially strong gains on boundaries and medium/small conditioned regions.

Xin Lin, Haodong Li, Zhi-Fei Zhang et al. · 1 citation
Open access Sep 2026

IP-ConTex: detail-consistent texture generation with image prompt

Textures are critical for enhancing the visual fidelity and diversity of three-dimensional (3D) models. Recently, generative models have significantly advanced texture generation. However, fine-grained control of the generation process remains challenging. Hence, we propose IP-ConTex, which is a novel image-guided text...

Lei Wang, Jie-Qing Feng · 0 citations
Conference Open access Sep 2026

Auxiliary text-guided image restoration for image-text matching

A new ITM framework that improves the model's discriminative performance by focusing on localized core attributes and has a better robustness in handling highly similar hard negatives, which provides a new way to address cross-modal hard sample discrimination.

Kuang-Rong Hao · 0 citations
Open access Aug 2026

Structured Creative Evaluation for Text-to-Image Generative AI Models

In recent years, text-to-image (T2I) generation models have made substantial progress, particularly in visual realism and the expression of prompt semantics. However, a key difficulty remains: how to evaluate generated results automatically in a way that is both comprehensive and interpretable, while still being practi...

W.-C. Ma, Q. Zhang · 0 citations
Conference Open access 2026

The Comprehensive Investigation of Controllable Image Generation Techniques

In recent years, the progress of diffusion models in text-to-image generation has been obvious to all. Text-based descriptions alone, however, are often insufficient for precisely controlling screen content. In order to solve this problem, researchers began to try to introduce additional conditions or visual references...

Hua-Man Fang · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.