PRAGMA is introduced, a benchmark for evaluating personalized guidance in long-term conversations and highlights the need for memory architectures that support robust conversational retrieval and memory-grounded reasoning beyond evidence recall.
Abstract
Large language models (LLMs) are increasingly deployed as personalized assistants that interact with users over extended periods of time. As conversations grow longer, relying on full interaction histories becomes increasingly inefficient and unreliable: long contexts introduce substantial computational overhead, making it difficult for models to consistently identify and utilize the most relevant information for the current request. These challenges have motivated memory systems that structure and retrieve user-specific information. In realistic interactions, users often seek practical guidance such as recommendations, planning, and decision support. Unlike factual recall tasks, personalized guidance requires models to integrate information across multiple past conversations and reason about changing user preferences and experiences. However, existing conversational memory evaluations mainly focus on retrieval and factual recall. To study this challenge, we introduce PRAGMA, a benchmark for evaluating personalized guidance in long-term conversations. PRGAMA contains curated longitudinal conversation histories, evidence annotations, and guidance scenarios grounded in evolving user contexts and incorrect user assumptions. Experiments across retrieval systems, memory systems, and long-context models reveal that current systems struggle both to recover the appropriate conversational evidence and to effectively use it for personalized guidance. Our results highlight the need for memory architectures that support robust conversational retrieval and memory-grounded reasoning beyond evidence recall.
PRISK is proposed, a dynamic evaluation framework with automated data generation and tailored metrics that uncovers systematic limitations in current LLM personalization and how personalized information shapes its responses.
Yumeng Wang, Yu-Chen Wu, Cheng Qian et al.· 0 citations
PersMem is proposed, a user profile memory framework for LLM personalization that addresses three key questions: what user information to store, how to organize it, and how to use it effectively during generation that consistently improves personalization effectiveness while reducing prompt length.
Yang-Xu Liao, Yong-Heng Deng, Tianyuan Jiang et al.· Proceedings of the 32nd ACM...· 0 citations
A structured memory framework for query-conditioned user-state inference for long-term personalization that achieves state-of-the-art performance on both PersonaMem and KnowU-Bench, demonstrating the effectiveness of query-conditioned user-state inference for long-term personalization.
Heng Wang, Yifei Li, Lingling Zhang et al.· 0 citations
It is hypothesize that existing benchmarks and user satisfaction are tracking different capabilities: benchmarks measure elicited retrieval (recall when asked), while conversation requires natural integration (detecting relevance and naturally weaving prior context into a response).
Ryuichi Sumida, K. Inoue, Tatsuya Kawahara· 0 citations
UtilMem is introduced, a diagnostic benchmark comprising 1,717 instances across five domains, designed to evaluate four underexplored aspects of memory utilization: reasoning over dense histories, identifying implicitly relevant memories, synthesizing distributed evidence into summaries, analyses, or plans, and resisti...
This work introduces ArborMem, an online memory framework that represents a long-running conversation as a navigable forest of interaction states that outperforms the strongest baselines on three established benchmarks and introduces BranchMemEval, a controlled diagnostic benchmark for interleaved and resumable interac...
Zongwei Lv, Yue-Meng Xu, Yilun Yao et al.· 0 citations
With $2.1 million funding from Google.org, the open-source Public Transit Intelligence Hub will unify public transit monitoring, operations, and passenger communication.
Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.