Skip to content

PRAGMA: Evaluating Personalized Guidance with Memory Alignment in Lifelong Conversations

Sep 2026 · 0 citations · 29 references
Computer Science

TL;DR

PRAGMA is introduced, a benchmark for evaluating personalized guidance in long-term conversations and highlights the need for memory architectures that support robust conversational retrieval and memory-grounded reasoning beyond evidence recall.

Abstract

Large language models (LLMs) are increasingly deployed as personalized assistants that interact with users over extended periods of time. As conversations grow longer, relying on full interaction histories becomes increasingly inefficient and unreliable: long contexts introduce substantial computational overhead, making it difficult for models to consistently identify and utilize the most relevant information for the current request. These challenges have motivated memory systems that structure and retrieve user-specific information. In realistic interactions, users often seek practical guidance such as recommendations, planning, and decision support. Unlike factual recall tasks, personalized guidance requires models to integrate information across multiple past conversations and reason about changing user preferences and experiences. However, existing conversational memory evaluations mainly focus on retrieval and factual recall. To study this challenge, we introduce PRAGMA, a benchmark for evaluating personalized guidance in long-term conversations. PRGAMA contains curated longitudinal conversation histories, evidence annotations, and guidance scenarios grounded in evolving user contexts and incorrect user assumptions. Experiments across retrieval systems, memory systems, and long-context models reveal that current systems struggle both to recover the appropriate conversational evidence and to effectively use it for personalized guidance. Our results highlight the need for memory architectures that support robust conversational retrieval and memory-grounded reasoning beyond evidence recall.

View source

Similar papers

Book Open access Aug 2026

Personalizing Large Language Models with User Profile Memory

PersMem is proposed, a user profile memory framework for LLM personalization that addresses three key questions: what user information to store, how to organize it, and how to use it effectively during generation that consistently improves personalization effectiveness while reducing prompt length.

Yang-Xu Liao, Yong-Heng Deng, Tianyuan Jiang et al. · 0 citations
Preprint Aug 2026

QUMem: Personalized Memory for Query-Conditioned User-State Inference in LLM Agents

A structured memory framework for query-conditioned user-state inference for long-term personalization that achieves state-of-the-art performance on both PersonaMem and KnowU-Bench, demonstrating the effectiveness of query-conditioned user-state inference for long-term personalization.

Heng Wang, Yifei Li, Lingling Zhang et al. · 0 citations
Preprint Aug 2026

MemUse: Moving Memory Evaluation from Direct QA to Natural Integration in Long-Term Human-AI Conversation

It is hypothesize that existing benchmarks and user satisfaction are tracking different capabilities: benchmarks measure elicited retrieval (recall when asked), while conversation requires natural integration (detecting relevance and naturally weaving prior context into a response).

Ryuichi Sumida, K. Inoue, Tatsuya Kawahara · 0 citations
#natural language process... Preprint Aug 2026

UTILMEM: Benchmarking Evidence Utilization in Long-Term Conversational Memory

UtilMem is introduced, a diagnostic benchmark comprising 1,717 instances across five domains, designed to evaluate four underexplored aspects of memory utilization: reasoning over dense histories, identifying implicitly relevant memories, synthesizing distributed evidence into summaries, analyses, or plans, and resisti...

Pei-Jun Qing, Fobo Shi, S. Vosoughi · 0 citations
Preprint Aug 2026

ArborMem: Navigating Interaction States with Memory Forests

This work introduces ArborMem, an online memory framework that represents a long-running conversation as a navigable forest of interaction states that outperforms the strongest baselines on three established benchmarks and introduces BranchMemEval, a controlled diagnostic benchmark for interleaved and resumable interac...

Zongwei Lv, Yue-Meng Xu, Yilun Yao et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.