MM-ContextFold: Context Folding for Multimodal Agentic Retrieval
MM-ContextFold is proposed, a training-free framework that loads raw images only when needed and maintains a persistent, text-only main context for high-level planning and spawns ephemeral branch contexts for image-dependent subtasks.
Yang Tian, Fan Liu, Jing-Yuan Zhang et al.
· 0 citations