MIRAGE: How Conversation State Shapes Historical Evidence Use in Multimodal Personal Agents
Multimodal large language model (MLLM) agents are increasingly used as personal assistants for long-running tasks. Their utility depends on continuity: agents must retrieve and use earlier evidence across dialogue, files, and workspace state. However, agents can generate plausible answers even when access to that histo...