Skip to content

SpeakerMem-R1: Speaker-Centered Dual-Track Memory for Multi-Party Dialogue

Sep 2026 · 0 citations · 88 references
Computer Science

TL;DR

SpeakerMem-R1 is proposed: its dual-track memory stores speaker-labeled verbatim messages and derived states organized into person-level and group-level views, then combines evidence from both tracks by entity, event, and time at query time, and ablations show that the verbatim and structured tracks, as well as person-level and group-level views, are complementary under the standardized evaluation interface.

Abstract

Long-term conversational memory in multi-party settings requires more than retrieving relevant content from long-term conversations: it must distinguish who said what, whom each statement concerns, how individuals perceive one another, what information is shared by the group, and how states change over time. Recent studies on multi-party dialogue benchmarks show that existing general-purpose LLM memory systems tend to lose person and group relations or struggle to integrate clues distributed across members, groups, and time. Together, these issues reveal two core bottlenecks: message attribution and relational understanding in multi-party dialogue, and state reconstruction from interleaved histories. To address both, we propose $\textbf{SpeakerMem-R1}$: its dual-track memory stores speaker-labeled verbatim messages and derived states organized into person-level and group-level views, then combines evidence from both tracks by entity, event, and time at query time. To reduce attribution and update errors during structured memory construction while enabling local deployment, we train Writer-R1 with SpeakerLevenshtein and speaker-conditioned GRPO. On GroupMemBench, SocialMemBench, and EverMemBench, SpeakerMem-R1 achieves binary accuracies of 47.9%, 69.2%, and 61.9%, respectively. On the publicly reported EverMemBench leaderboard from EverMind-AI, we achieves 62.33%, the best reported result among the latest state-of-the-art frameworks. It also achieves 70.85% on all 1,986 LoCoMo questions, which we use as a two-person long-term conversation boundary test. In a controlled evaluation of 305 questions, RL raises the SFT Writer's mean accuracy from 57.38% to 68.20%. We report both binary accuracy and token-F1, and ablations show that the verbatim and structured tracks, as well as person-level and group-level views, are complementary under the standardized evaluation interface.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

Beyond Dyadic Memory: Interaction-Aware Multimodal Memory with Adaptive Agentic Retrieval for Multi-Party Spoken Conversations

Long-term memory enables agents to accumulate information and reason across sessions, yet existing research primarily focuses on dyadic text or image-text conversations, leaving long-term memory for multi-party spoken conversations underexplored. This setting requires preserving conversational content, identifying part...

Wen-Xu Jia, Xi-Ze Cheng, Zi-Han Zhang et al. · 0 citations

TEIDAN: A Multilingual Multiparty Dialogue Corpus

This paper describes the collection design, participants, recording setup, transcription format, and corpus statistics, and provides preliminary analyses to illustrate how TEIDAN can support research on turn-taking, addressee recognition, and multimodal grounding in human-human and human-agent interaction.

Taiga Mori, K. Inoue, Mikey Elmers et al. · 1 citation
#artificial intelligence Preprint Sep 2026

Who Said What, and Will It Be Remembered? Evaluating Persistent Speaker Attribution Across Meetings

Speech transcripts used as long-term memory must preserve both words and stable speaker identities. Existing meeting-transcription metrics either ignore speakers or remap anonymous speakers independently in each recording, so they cannot measure whether the same person retains one identity across meetings. We evaluate...

Shantanu Vispute, A. M. Mishra, Siddhartha Saxena · 0 citations
Preprint Sep 2026

Voices as Handles: Reasoning about Speaker Identity with Frozen Text LLMs

Multi-user voice agents must track who said what across dialogue sessions. Text LLMs are attractive backbones for such agents, but transcripts alone do not expose acoustic speaker identity, leaving the model without a persistent reference for linking information to speakers across sessions. We address this gap by intro...

Run-Qiu Xu, Zhi-Sheng Zheng, David Harwath · 0 citations
Preprint Aug 2026

FTA-Mem: Fact-Time-Affect Anchored Memory for Low-Density Long-Term Dialogue

Long-term emotional-support agents require memory mechanisms for personalized understanding across sessions. However, emotional-support dialogue is often low-density: turns are incomplete, evidence is scattered, and user states evolve over time. Existing memory methods usually rely on fixed units, such as turn-level no...

Chang Liu, Shu-Yi Zhang, Chang-Sheng Ma et al. · 0 citations
Preprint Aug 2026

MemUse: Moving Memory Evaluation from Direct QA to Natural Integration in Long-Term Human-AI Conversation

It is hypothesize that existing benchmarks and user satisfaction are tracking different capabilities: benchmarks measure elicited retrieval (recall when asked), while conversation requires natural integration (detecting relevance and naturally weaving prior context into a response).

Ryuichi Sumida, K. Inoue, Tatsuya Kawahara · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

MIT News · Artificial Intelligence Aug 27, 2026

Looking beyond natural sequences

A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.