Skip to content

Where to Look and What to Use: Retrieve-Localize-Generate for Long-Term Conversational Memory Question Answering

Sep 2026 · 0 citations · 35 references
Computer Science

TL;DR

This work proposes MemLoc, a unified Retrieve-Localize-Generate framework for long-term conversational memory QA, and introduces a reasoning-based evidence locator trained with Self-reflective Hint Policy Optimization, which performs progressive refinement by extracting query-relevant fragments within memory units to suppress noise.

Abstract

Retrieval-augmented generation (RAG) enables large language models (LLMs) to answer questions by accessing external knowledge and has been widely adopted for long-term conversational memory question answering. However, existing methods suffer from two key challenges: (1) fragmented evidence scattered across temporally distant sessions, and (2) noisy content within retrieved sessions that triggers the lost-in-the-middle effect. To address these challenges, we propose MemLoc, a unified Retrieve-Localize-Generate framework for long-term conversational memory QA. For retrieval, MemLoc decomposes each session into multi-granularity memory units and performs query routing via an inner-memory graph with entropy-based granularity selection. It further models cross-session semantic and temporal dependencies through a cross-memory graph, enabling coarse-to-fine retrieval of top-K relevant memory candidates. For localization, we introduce a reasoning-based evidence locator trained with Self-reflective Hint Policy Optimization (SHPO), which performs progressive refinement by extracting query-relevant fragments within memory units to suppress noise and reranking across candidates to remove redundancy, producing a compact evidence set with lightweight location IDs. For generation, these IDs act as precise grounding signals that guide the LLM to the correct memory positions, mitigating the lost-in-the-middle effect while preserving original contextual integrity. Extensive experiments on four benchmarks demonstrate that MemLoc achieves state-of-the-art retrieval accuracy and response quality while maintaining efficiency. Our code is available at: https://github.com/Nikol-coder/MemLoc.

View source

Similar papers

#natural language process... Preprint Sep 2026

JustMem: Just-Enough Memory Access for Long-Term Conversations

JustMem is introduced, which stores conversation history as compact atomic memories and adapts memory access along two dimensions to each query and achieves the highest mean accuracy and retrieval recall among the compared memory systems while using substantially fewer generative-model tokens for memory construction an...

Guan-Hua Chen, Yan-Ting Wang, Wen-Jing Zhi et al. · 1 citation
Book Open access Aug 2026

MCoRe: Multi-Entry Complementary Retrieval with Reflection-Guided Iteration for Multi-Hop QA

MCoRe, a multi-entry complementary retrieval framework with reflection-guided iteration for multi-hop QA that enables multi-entry complementary retrieval by indexing entry units at multiple semantic resolutions with explicit links to chunk evidence, and fusing cross-resolution hits via chunk-level voting to form a comp...

Ju-Xiang Zeng, Zhuohui Gao, Zhe Hou et al. · 0 citations
Preprint Aug 2026

EnSI-RAG: Entity-Structure-Indexed Retrieval-Augmented Generation for Long-Document Question Answering

This work proposes EnSI-RAG (Entity-Structure-Indexed Retrieval-Augmented Generation), a framework that constructs a query-independent, entity-centered index that separates evidence localization from answer synthesis while preserving traceable source evidence.

Xuan-Yu Meng, Jiashuo Sun, Jash Parekh et al. · 1 citation
#artificial intelligence Preprint Sep 2026

Remember by Asking: Retrieval-Induced Memory Evolution for LLM Agents

RIME is introduced, a retrieval-induced memory framework that shifts memory construction from monolithic compression toward evidence-centered integration and consistently achieves the best performance across all three quality metrics among the compared methods, while requiring substantially fewer query-time LLM tokens.

Wan-Qi Zhou, Jia-Wei Lu, Yang Wang et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.