Long-horizon embodied interaction requires agents to retain and continually update information about the environment as they observe, act, and encounter change. Yet current agents struggle to maintain such memory reliably. Our analysis traces this limitation to four key deficiencies: weak fine-grained visual memory, un...
Li-Zhou Liang, Xin-Yu Zhong, Miao Pan et al.· 0 citations
Effective memory is crucial for LLM agents, yet constructing it effectively remains challenging. A memory-construction policy decides what information to extract, store, update, compress, or discard as interactions accumulate. Heuristic memory methods rely on subjective, task-specific rules, which can misalign with dow...
TASPO, which converts privileged supervision into outcome-grounded action credit, is introduced and indicates that TASPO reduces supervision mismatch and that action-level assignment stabilizes the policy optimization process.
Jing-Xiao Yang, Wang-Jie Gan, Ying-Xuan Zhuang et al.· 4 citations
Experiments show that AttriMem outperforms retrieval-based, heuristic, and RL-based baselines, generalizes across benchmarks and answer models, stabilizes RL optimization, and outperforms retrieval-based, heuristic, and RL-based baselines on long-horizon dialogue question answering.
Qin-Feng Li, Yun-Tai Bao, Xinyang Yu et al.· 0 citations
SkillAligner is proposed, a training-free execution-time skill adaptation framework that treats retrieved skills as adaptable drafts rather than fixed instructions that substantially improves task performance over existing skill-use baselines, reduces skill-induced regressions at the instance level, and lowers total in...
Qin-Feng Li, Dalin He, Yun-Tai Bao et al.· 0 citations
This work proposes TAMP-Nav, a unified framework for efficient embodied navigation that dynamically triggers Chain-of-Thought and retains high-fidelity memory only at critical nodes, compressing redundant trajectories into lightweight Space-Time Indicators, thereby preserving critical historical information and enhanci...
Hongyan Feng, Sun-Lai Chen, Xuan-Yu Liu et al.· 1 citation
ProVisE (Protocolized Visual Evaluation), a benchmark-agnostic framework that elicits protocol-constrained visual answers from image-generation models and parses them into structured predictions compatible with original metrics, is proposed and revealed, revealing complementary strengths of pixel-space expression and t...
Xu Wang, Kaixiang Yao, Miao Pan et al.· arXiv.org· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.