Multimodal large language model (MLLM) agents are increasingly used as personal assistants for long-running tasks. Their utility depends on continuity: agents must retrieve and use earlier evidence across dialogue, files, and workspace state. However, agents can generate plausible answers even when access to that histo...
Yu Liu, Wen-Xiao Zhang, Cheng Hu et al.· 0 citations
CARE (Canonicalization, Attribution, and Resolution Engine), a shell-specific, static-first verifier for individual shell commands before execution can reduce dispatch-boundary risk for LLM agents while preserving most benign workflows.
Yu Liu, Wenxiao Zhang, Zhiwei Yang et al.· arXiv.org· 1 citation
Experiments across multiple long-context narrative question answering and claim verification settings show that ClueWeaver substantially improves local end-to-end language models while providing evidence coverage and paragraph-referenced reasoning traces.
Ji-Hao Zhu, Zhi-Wei Yang, Wen-Xiao Zhang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.