Skip to content
Preprint

MemUse: Moving Memory Evaluation from Direct QA to Natural Integration in Long-Term Human-AI Conversation

Aug 2026 · 0 citations · 30 references
Computer Science

TL;DR

It is hypothesize that existing benchmarks and user satisfaction are tracking different capabilities: benchmarks measure elicited retrieval (recall when asked), while conversation requires natural integration (detecting relevance and naturally weaving prior context into a response).

Abstract

Memory systems for conversational LLMs are conventionally evaluated by direct, fact-seeking questions about prior dialogue (Direct QA): can the model recall fact X from a prior conversation? We tested whether higher Direct QA accuracy correlates with higher user satisfaction in a 4-month deployment (40 users, 1,872 sessions, 7 memory conditions). Existing-benchmark Direct QA varies from 19.7% to 70.1% across the 7 conditions, but satisfaction does not change. We hypothesize that existing benchmarks and user satisfaction are tracking different capabilities: benchmarks measure elicited retrieval (recall when asked), while conversation requires natural integration (detecting relevance and naturally weaving prior context into a response). To examine this, we introduce MemUse, a set of real user-cued memory moments drawn from the deployment, scored by an integration-aware judgment of the natural conversational response. Holding the model and context fixed, the same system that scores 78.8% on Direct QA references only 7.9% of those facts in conversation -- a 71-point gap. Within these moments, Natural Integration is associated with satisfaction, whereas Direct QA is not. We release the deployment corpus and MemUse together with all judgments and scoring prompts at https://github.com/ryuichi-sumida/memuse.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

When Users Don't Ask: Benchmarking Context-Driven Memory Retrieval in Conversational Agents

Large language models (LLMs) are increas- ingly deployed as long-horizon conversational agents, motivating growing interest in mem- ory systems. However, existing benchmarks primarily evaluate memory through QA-style probing rather than in-situ conversational usage. We introduce LOCOMO-CONV, a conversa- tional memory b...

Wen-Yu Chang, Yun-Nung Chen · 3 citations · ⚡1
Preprint Aug 2026

FTA-Mem: Fact-Time-Affect Anchored Memory for Low-Density Long-Term Dialogue

Long-term emotional-support agents require memory mechanisms for personalized understanding across sessions. However, emotional-support dialogue is often low-density: turns are incomplete, evidence is scattered, and user states evolve over time. Existing memory methods usually rely on fixed units, such as turn-level no...

Chang Liu, Shu-Yi Zhang, Chang-Sheng Ma et al. · 0 citations
#natural language process... Preprint Sep 2026

RuleMem: Active Rule Memory for Long-Term Conversational Agents

A rule-based memory framework that induces reusable logical rules from historical interactions to guide both evidence retrieval and reasoning, and constructs natural-language Horn clauses from conversations and validates them via a Rule Perplexity Consistency (RPC) mechanism.

Xing-Yuan Zeng, Zuo-Han Wu, Quanming Yao et al. · 0 citations
#natural language process... Preprint Sep 2026

CueMem: Cue-Guided Context Reconstruction for Long-Term Conversational Memory

Long-term conversational agents must answer user queries by recalling information from extended dialogue histories, yet directly using the full history is costly and often unreliable, while compressed memory units may lose fine-grained evidence needed for question answering. Motivated by the reconstructive view of auto...

Chang-Jian Wang, Rong-Zhen Li, Wei-Li Guan et al. · 0 citations
Book Open access Sep 2026

Consistent Conversational State for Virtual Agents: Slot-Based Memory for Accurate Fact Retrieval

This work investigates whether memory interference originates mainly from memory retrieval or from the accumulation of competing fact versions added during memory updates, and evaluates how memory-write policies influence memory retrieval behavior later on.

Erica Butts, Salam Daher · 0 citations
#natural language process... Preprint Aug 2026

UTILMEM: Benchmarking Evidence Utilization in Long-Term Conversational Memory

UtilMem is introduced, a diagnostic benchmark comprising 1,717 instances across five domains, designed to evaluate four underexplored aspects of memory utilization: reasoning over dense histories, identifying implicitly relevant memories, synthesizing distributed evidence into summaries, analyses, or plans, and resisti...

Pei-Jun Qing, Fobo Shi, S. Vosoughi · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.