MM-ContextFold is proposed, a training-free framework that loads raw images only when needed and maintains a persistent, text-only main context for high-level planning and spawns ephemeral branch contexts for image-dependent subtasks.
Abstract
Multimodal Agentic Retrieval (MAR) requires agents to solve complex information-seeking tasks by iteratively invoking external tools. Typical frameworks such as ReAct maintain raw multimodal inputs and the accumulating interaction history in a single, ever-growing context, leading to the context explosion problem. While existing methods alleviate this issue by compressing redundant text, effective strategies for managing token-intensive visual content remain largely underexplored. To address this gap, we first conduct a systematic empirical study of approximately 10,000 trajectories. The results show that as visual cues are progressively extracted through external tools and textualized into the context, raw images become increasingly redundant. Continued image retention is associated with higher output entropy and can even degrade task accuracy. Motivated by these findings, we propose MM-ContextFold, a training-free framework that loads raw images only when needed. It maintains a persistent, text-only main context for high-level planning and spawns ephemeral branch contexts for image-dependent subtasks. Within each branch, the agent loads the relevant images, completes the subtask, and folds the result back into the main context as a concise textual summary; the images and branch trace are then discarded. Experiments on seven MAR benchmarks across five backbone models show that MM-ContextFold improves average accuracy by 6.3 percentage points over ReAct while reducing the working context length by 27.5\%.
WeAgent-Harness, a multimodal agentic harness that supports native text-vision interaction and runtime recovery, and WeAgent-MMSearch, an integrated system spanning data construction, agentic post-training, and multimodal rollout that outperform similarly sized open-source models and rival models with roughly ten times...
Zongkai Liu, Hui Zhang, Li-Qiang Niu et al.· 0 citations
A new model based on a Large Multimodal Model (LMM) that functions as a context-aware reasoner for CoCo-IR is proposed, which interprets the entire interaction history to generate Transformable Image Embeddings (TIE) that evolve across turns.
Shengcao Cao, T. Dabral, Z. Ding et al.· 0 citations
This work presents WeAgent-MMGenEdit, a full-stack recipe including a multimodal harness, a scalable data construction pipeline, a comprehensive benchmark, and post-training methods for the agent policy and image backend that enables a 30B-total/3B-active policy to outperform similarly sized policy models and approach...
Hui Zhang, Zongkai Liu, Li-Qiang Niu et al.· 0 citations
This work proposes a novel Human-profile Enhanced Retrieval Optimization framework for long-term agent memory (HERO), which converts the dialogue history into a traceable heterogeneous memory graph that preserves raw dialogue text as evidence for reasoning, thereby mitigating information loss.
Yuanhua Lin, Yile Li, Zhiyuan Zhao et al.· 0 citations
Across multi-step visual tool-use benchmarks, SPARE achieves the highest average accuracy among pruning methods while removing 37.89-64.58\% of reasoning tokens, showing that reducing textual dominance restores reliance on visual evidence and mitigates over-conditioning on self-generated language.
Yuchen Huang, Sijia Li, Jun Zhang et al.· 0 citations
Retrieval-Augmented Generation (RAG) enhances Large Language Models (LLMs) by reducing hallucinations and improving factual accuracy. Most RAG implementations, however, are restricted to a single data modality. Real-world document collections span PDFs, spreadsheets, audio, video and images simultaneously. This paper i...
S. Vijayan, Sushila Palwe, Devendra Joshi· International Conference on...· 0 citations
Known for his clear and elegant writing style, Bertsekas shaped fields from control and optimization to large-scale computation and artificial intelligence.