Skip to content

Propose, Verify, Commit: Evidence-Grounded Memory for Long-Horizon Multi-Actor Conversations

Sep 2026 · 0 citations · 16 references
Computer Science

TL;DR

This work introduces EGMEMORY, which formulates long-horizon multi-actor memory as a searchable state machine that separates persistent message-level evidence from an explicit active state.

Abstract

Long-horizon conversational memory is especially challenging in multi-actor settings, where relevant evidence is distributed across participants and contexts and previously established information may later be revised. We introduce EGMEMORY, which formulates long-horizon multi-actor memory as a searchable state machine that separates persistent message-level evidence from an explicit active state. At write time, adaptive state resolution and an evidence-grounded propose-verify-commit protocol govern how this state evolves. At read time, adaptive evidence navigation iteratively resolves the state and supporting evidence required for a query, using conversational structure to narrow the search space and lexical-semantic relevance to rank candidates. The system operates through prompting and tool use without memory-specific policy training. EGMEMORY achieves 68.2% on GroupMemBench and 77.9% on EverMemBench, outperforming the strongest evaluated baselines by 22.7 and 21.4 percentage points, respectively. It further reaches 73.6% on the dyadic LoCoMo benchmark, demonstrating generalization beyond multi-actor conversations. We will release the codebase upon formal publication.

View source

Similar papers

Preprint Aug 2026

Reconstructing the Right Episode: Evaluating Interleaved Conversational Memory Beyond Long Context

Conversations with chat assistants increasingly span many topics in a single long-running thread, challenging memory systems. Existing long-context and memory benchmarks often expose session or topic boundaries, or probe direct personal-memory questions. These settings understate a harder assistant-memory regime: a fla...

Zhe-Xi Feng, Rui-Yi Zhang, Yong-Bo Yang et al. · 0 citations
#natural language process... Preprint Sep 2026

Beyond Memory Construction: Rethinking Memory Access for LLM-based Conversational Agents

Memory is a core component of conversational agents, enabling coherent and context-aware behavior over long interactions. Recent approaches commonly rely on LLM-based memory construction, where raw interactions are rewritten into structured memory units and later retrieved via a RAG pipeline. While effective in control...

Dong-Hua Cai, Yong-Heng Deng, Yi-Fei Wang et al. · 0 citations
Preprint Aug 2026

Can Agent Memory Systems Track Evolving State?

StateMem is presented, a state-first memory method that explicitly tracks supersession and relational dependencies, and it is shown it improves current-state accuracy over the strongest same-backbone baseline and over the strongest memory system, while remaining competitive with the long-context baselines.

Xin-Yi Fan, Miri Liu, Ruozhen Yang et al. · 4 citations
#natural language process... Preprint Aug 2026

UTILMEM: Benchmarking Evidence Utilization in Long-Term Conversational Memory

UtilMem is introduced, a diagnostic benchmark comprising 1,717 instances across five domains, designed to evaluate four underexplored aspects of memory utilization: reasoning over dense histories, identifying implicitly relevant memories, synthesizing distributed evidence into summaries, analyses, or plans, and resisti...

Pei-Jun Qing, Fobo Shi, S. Vosoughi · 0 citations
Preprint Aug 2026

MemPrism: Task-Conditioned Relational Memory Views for Long-Horizon Agents

Experiments on long-horizon embodied and web-agent benchmarks show that MemPrism consistently improves the task performance, especially as trajectories become longer, while reducing memory token consumption.

Zhi-Sheng Chen, Bingfan Zeng, Bangde Cao et al. · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.