Skip to content

Agent-Editing World Model: Rethinking World Modeling for LLM Agents

Sep 2026 · 0 citations
Computer Science

TL;DR

This work proposes the Agent-Editing World Model (AEWM), which models how reasoning and actions shape future task progress rather than simulating tool responses, and combines AEWM with EditAct, which integrates these capabilities with real execution.

Abstract

Recent advances in large language models (LLMs) have enabled agents to tackle long-horizon tasks across diverse environments. To further improve agent performance, existing language world models typically predict environment observations, yet reconstructing high-entropy, execution-dependent tool responses offers limited value when real feedback is available. Meanwhile, agents suffer from \emph{task-state contamination}, where unsupported assumptions and outdated plans persist in history and distort subsequent decisions. We propose the \textbf{Agent-Editing World Model (AEWM)}, which models how reasoning and actions shape future task progress rather than simulating tool responses. AEWM combines \textbf{Action Judge} to distinguish \textsc{Critical}, \textsc{Exploratory}, and \textsc{Noisy} decisions with \textbf{State Revision} to edit noisy reasoning--action continuations from the same observed history. \textbf{EditAct} integrates these capabilities with real execution, directly changing the state underlying subsequent decisions rather than merely providing critiques. We train AEWM across Search, Terminal, and Software Engineering through mid-training and supervised fine-tuning. AEWM achieves 70.5\% macro-F1 on our Action Judge benchmark, exceeding the strongest frontier baseline by 10.6 points. Across six benchmarks and three agent backbones, EditAct improves average scores by 3.2--6.7 points over the strongest baseline. Furthermore, rejection sampling fine-tuning on verified EditAct trajectories, termed \textbf{AEWM-RFT}, improves over Self-RFT by 2.2--2.6 points across three domains without online AEWM guidance.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

Rep2Skill: Representation-Guided Skill Self-Evolution for LLM Agents

Textual skills enable large language model (LLM) based agents to accumulate reusable procedural knowledge without updating model parameters. Yet existing skill evolution remains largely confined to the text space: an optimizer must diagnose success and failure patterns, and revise skills solely from long execution traj...

Kai-Xin Zhang, Chang-Ming Li, Ying-Dong Shi et al. · 0 citations
#artificial intelligence Preprint Sep 2026

When Can Agents Forget Their Reasoning? ICLR for Long-Horizon Agent Context Compression

Long horizon language model agents continually accumulate reasoning history, increasing context length and inference cost even after earlier decisions have been executed and observed. Unlike static Chain of Thought compression, removing historical reasoning can change future actions and the resulting interaction trajec...

Ming-Xuan Wang, Fei Luo, Bo Wang et al. · 0 citations
Preprint Aug 2026

Reading is not Reasoning: Bridging the Agentic Policy Gap in Vision-Text Compression

Multi-step language-model agents repeatedly process growing interaction histories, leading to substantial context costs. Vision--text compression reduces these costs by rendering history as images, but the resulting modality shift creates a marked capability gap. Through controlled evaluations of history recovery, matc...

Cheng Fan, Junyi Zhou, Tingzhang Luo et al. · 1 citation
#artificial intelligence Preprint Sep 2026

AgentBrew: Offline Tool-Use Agent Learning from Raw Real-World Trajectories

AgentBrew is proposed, an offline training framework that learns effective tool-use policies from a single batch of raw interaction trajectories, without task verifiers or iterative on-policy rollouts, and demonstrates that fine-grained offline learning can recover useful supervision from raw trajectories that filterin...

Zhiyi Lyu, Ye-Wen Li, Long-Tao Zheng et al. · 2 citations
Preprint Sep 2026

Code Plans, Diffusion Renders: Open-Ended Generative World Modeling

Extensive experiments demonstrate that the CoDeR framework substantially extends the capabilities of existing world models, enabling long-term memory, open-ended interactions, autonomous evolution, and persistent multi-agent dynamics, while achieving state-of-the-art performance across multiple evaluation settings.

Zi-Xun Fang, Ya-Wen Shao, Kai Zhu et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

MIT News · Artificial Intelligence Aug 27, 2026

Looking beyond natural sequences

A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.