Long-horizon embodied interaction requires agents to retain and continually update information about the environment as they observe, act, and encounter change. Yet current agents struggle to maintain such memory reliably. Our analysis traces this limitation to four key deficiencies: weak fine-grained visual memory, un...
Li-Zhou Liang, Xin-Yu Zhong, Miao Pan et al.· 0 citations
Effective memory is crucial for LLM agents, yet constructing it effectively remains challenging. A memory-construction policy decides what information to extract, store, update, compress, or discard as interactions accumulate. Heuristic memory methods rely on subjective, task-specific rules, which can misalign with dow...
TASPO, which converts privileged supervision into outcome-grounded action credit, is introduced and indicates that TASPO reduces supervision mismatch and that action-level assignment stabilizes the policy optimization process.
Jing-Xiao Yang, Wang-Jie Gan, Ying-Xuan Zhuang et al.· 4 citations
Experiments show that AttriMem outperforms retrieval-based, heuristic, and RL-based baselines, generalizes across benchmarks and answer models, stabilizes RL optimization, and outperforms retrieval-based, heuristic, and RL-based baselines on long-horizon dialogue question answering.
Qin-Feng Li, Yun-Tai Bao, Xinyang Yu et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.