Effective memory is crucial for LLM agents, yet constructing it effectively remains challenging. A memory-construction policy decides what information to extract, store, update, compress, or discard as interactions accumulate. Heuristic memory methods rely on subjective, task-specific rules, which can misalign with dow...
Experiments show that AttriMem outperforms retrieval-based, heuristic, and RL-based baselines, generalizes across benchmarks and answer models, stabilizes RL optimization, and outperforms retrieval-based, heuristic, and RL-based baselines on long-horizon dialogue question answering.
Qin-Feng Li, Yun-Tai Bao, Xinyang Yu et al.· 0 citations
A rationale-guided knowledge distillation framework for cross-lingual stance detection using Chain-of-Thought prompting to guide Large Language Models in generating informative rationales, and distill the resulting reasoning knowledge into a compact student model.