Skip to content

AcFlow: Controlling Text-to-Image Diffusion Transformers via Learned Conditional Activation Flow

Sep 2026 · 0 citations · 51 references
Computer Science

TL;DR

AcFlow is introduced, an inference-time controller that transports intermediate layer image-token activations through a learned concept-conditioned velocity field while keeping the base DiT frozen, and supports the learned velocity field as an adaptive control mechanism.

Abstract

Text-to-image diffusion transformers (DiTs) are powerful generators, yet direct prompting provides limited control interface for style intensity and can fail to suppress unwanted concepts. To enable these controls, we introduce AcFlow, an inference-time controller that transports intermediate layer image-token activations through a learned concept-conditioned velocity field while keeping the base DiT frozen. A textual concept description specifies the desired intervention, while the integration horizon provides a continuous control parameter. The field produces token-varying, activation-dependent updates. With parameters shared across concepts within each task family, one field covers over 15,000 style descriptions or over 1,000 suppression concepts, and generalizes to concepts unseen during training without per-concept fitting. On style control, AcFlow achieves the best style--content trade-off among the evaluated baselines in the high-style-alignment regime. At a fixed operating point, AcFlow attains style--content alignment of 0.5365/0.2860, compared with 0.4397/0.2684 for the baseline with the highest style alignment. On concept suppression, AcFlow reduces the fraction of images showing the concept from 95.3%/82.1% to 41.6%/40.5% on held-in/held-out concepts, including cases where deleting them from the prompt fails to remove them. Our analyses support the learned velocity field as an adaptive control mechanism, with update directions varying across tokens and depending on their activation states. Our code is available at https://github.com/Nove1yst/AcFlow.

View source

Similar papers

Preprint Aug 2026

HPSD: Hybrid-Policy Self-Distillation for Text-Image-to-Video Diffusion Models

This work proposes Hybrid-Policy Self-Distillation (HPSD), a novel self-distillation framework where a single TI2V model acts as both teacher and student under different conditions: the teacher operates in TI2V mode with a high-quality first frame and an enhanced prompt, while the student runs in the base T2V mode with...

Jia-Zi Bu, Peng-Yang Ling, Yu-Jie Zhou et al. · 2 citations

FlowLess: Controlling Abstract Image Generation

A novel self-supervised framework that enables granular control over image generation through a visual abstraction set that provides a richer, more flexible paradigm for creative design compared to state-of-the-art baselines across diverse styles and compositions is introduced.

Amir Hertz, Noah Snavely, Google DeepMind · 0 citations
#artificial intelligence Preprint Sep 2026

V-Engram: Trigger-Indexed External Memory for Modular Text-to-Image Personalization

Pretrained text-to-image models contain broad visual knowledge, yet they cannot reliably acquire or refine a specific visual identity from only a few references while preserving compositional control. Token-embedding methods are compact but often underfit identity, whereas adapter-based methods improve fidelity through...

Hao-Ran He, Run-Yuan Cai, Yi-Ming Wang et al. · 0 citations
Preprint Aug 2026

Semantic Steering for Controllable Generation: Tuning-Free Concept Erasure in Multimodal Diffusion Transformers

This work proposes to erase concepts by directly manipulating the model's internal representations by operating exclusively on the sparse text-branch tokens and leveraging the straight sampling trajectory of rectified flow, achieving effective concept erasure with negligible overhead and without any training.

Qiao Li, Xiaomeng Fu, Yuanshu Zhao et al. · 1 citation
Preprint Aug 2026

Diffusion Image Editing via Asynchronous Token Decoding

This approach combines local editing and background preservation without external or user-provided spatial masks and without model fine-tuning, and achieves the strongest reported preservation metrics, including 27.44~dB PSNR and 0.055 LPIPS.

Yang Shi, Liangsi Lu, Minzhe Guo et al. · 0 citations
Conference Aug 2026

Noun Presence Loss with Dynamic Weighting for Compositional Text-to-Image Generation

Compositional text-to-image generation requires faithful depiction of multiple objects with distinct visual attributes. While recent inference-time optimization methods have substantially improved attribute binding, entity neglect - where specified objects are absent from the generated image - remains an unresolved fai...

Trong-Tai Dam Vu, Vinh-Tiep Nguyen · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.