Skip to content
Preprint

WeAgent-MMGenEdit: A Full-Stack Recipe for Multimodal Agentic Image Generation and Editing

Sep 2026 · 0 citations · 100 references
Computer Science

TL;DR

This work presents WeAgent-MMGenEdit, a full-stack recipe including a multimodal harness, a scalable data construction pipeline, a comprehensive benchmark, and post-training methods for the agent policy and image backend that enables a 30B-total/3B-active policy to outperform similarly sized policy models and approach the performance of a 1T-parameter agent.

Abstract

Image generation and editing models have advanced rapidly, yet remain unreliable when prompts require external world knowledge. Bounded and long-tail parametric knowledge prevents direct or reason-then-generate approaches from recovering the required facts and visual appearances. Existing agentic generation and editing methods mitigate this limitation with retrieval tools, yet remain constrained by insufficient visual verification, overloaded policy models, and weak integration of retrieved textual and visual evidence. To address these limitations, we present WeAgent-MMGenEdit, a full-stack recipe including a multimodal harness, a scalable data construction pipeline, a comprehensive benchmark, and post-training methods for the agent policy and image backend. We first introduce WeAgent-Harness, a multimodal runtime with persistent evidence management and dedicated verification and integration tools that organize retrieved multimodal evidence into a dense carrier. Upon this, we develop a scalable pipeline for prompt synthesis and agentic trajectory collection, yielding 23K supervised trajectories and 14.7K RL tasks with three-layer verifiable checklists. We further introduce WeBench-MMGenEdit, a bilingual benchmark covering both knowledge-intensive image generation and multi-image editing. Finally, a two-sided post-training recipe based on SFT and RL improves the agent policy and image backend. Together, WeAgent-MMGenEdit enables a 30B-total/3B-active policy to outperform similarly sized policy models and approach the performance of a 1T-parameter agent.

View source

Similar papers

#artificial intelligence Preprint Aug 2026

WeAgent-MMSearch: Native Text-Vision Interaction for Multimodal Search Agents

WeAgent-Harness, a multimodal agentic harness that supports native text-vision interaction and runtime recovery, and WeAgent-MMSearch, an integrated system spanning data construction, agentic post-training, and multimodal rollout that outperform similarly sized open-source models and rival models with roughly ten times...

Zongkai Liu, Hui Zhang, Li-Qiang Niu et al. · 0 citations
Preprint Aug 2026

ToolArtist: Tool-Using Unified Multimodal Models for Agentic Image Generation

Text-to-image (T2I) models can produce visually compelling images, yet they remain limited on open-world tasks that require complex semantic understanding, multi-step reasoning, and the integration of external world knowledge. Existing efforts introduce agent capabilities into image generation, but they either prescrib...

Jiahao Zhao, Xiao-Min Yu, ZhongXiang Sun et al. · 3 citations
Preprint Aug 2026

Beyond Trial-and-Error: Agentic Optimization for Image-to-Video Adherence

Modern black-box Image-to-Video (I2V) models offer powerful capabilities in automated content creation, yet their lack of fine-grained control and reliability presents significant challenges in professional workflows. Their inherent stochasticity causes minor variations in textual prompts or hyperparameters to yield dr...

Aman Tyagi, Hemanth Boinpally, Jonathan Chen et al. · 1 citation
Preprint Aug 2026

Exploring the Performance Frontier of Compact Unified Image Generation Models

Swift-Image achieves leading aggregate performance among evaluated open-source models with only 6B parameters and 243K GPU training hours; the compressed 3B model incurs nearly no loss, while few-step distillation further improves aggregate editing performance with substantially fewer sampling steps.

Taihang Hu, Zhaowen Wang, Zuan Gao et al. · 0 citations
#artificial intelligence Preprint Sep 2026

MM-ContextFold: Context Folding for Multimodal Agentic Retrieval

Multimodal Agentic Retrieval (MAR) requires agents to solve complex information-seeking tasks by iteratively invoking external tools. Typical frameworks such as ReAct maintain raw multimodal inputs and the accumulating interaction history in a single, ever-growing context, leading to the context explosion problem. Whil...

Yang Tian, Fan Liu, Jing-Yuan Zhang et al. · 0 citations
Preprint Aug 2026

Open Evaluation Agent: Efficient and Promptable Evaluation of Visual Generative Models

The Evaluation Agent framework is proposed, which employs human-like strategies for efficient, dynamic, multi-round evaluations, offering detailed, user-tailored analyses and is efficient, promptable, explainable, and scalable across models and tools.

Shu-Lin Tian, Zi-Qi Huang, Fan Zhang et al. · 2 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.