Skip to content
Open access

Navigating State Drift in Retrieval-Augmented Generation (RAG) Agents: A Diagnostic Benchmark and Structured Graph-Guided Repair Framework

Sep 2026 · International Journal of Artificial Intelligence and Agent Systems · Vol 1 · 0 citations · 21 references

TL;DR

This work lays the foundation for building more reliable, state-aware RAG agents that resist drift across complex multi-turn interactions and introduces a comprehensive taxonomy covering goal, constraint, evidence, retrieval, tool-state, memory, and abstention drift.

Abstract

Multi-turn Retrieval-Augmented Generation (RAG) agents are increasingly relied upon to integrate conversational history, external knowledge retrieval, tool use, and persistent memory. However, their failures extend beyond isolated hallucinations to a more systemic and gradual degradation we term state drift, a phenomenon where the agent progressively loses alignment with the user's current goal, active constraints, retrieved evidence, tool observations, or updated memory. This paper offers three contributions. First, we formalize state drift and introduce a comprehensive taxonomy covering goal, constraint, evidence, retrieval, tool-state, memory, and abstention drift. Second, we present DRIFTBENCH, a diagnostic benchmark featuring controlled state perturbations and fine-grained state-level annotations to systematically elicit and measure drift in RAG agents. Third, we propose SAGE-R, a model-agnostic state-graph framework that diagnoses drift in real time and triggers typed repair actions including rollback, re-retrieval, tool cross-validation, clarification, and abstention. The current repository provides a deterministic v0 scaffold with generated examples, baseline proxies, repair traces, analysis tables, and a human-validation packet, establishing a validated research pipeline prior to the integration of external datasets, model-backed experiments, and completed human annotation. All quantitative results reported in this paper are therefore scaffold-level pipeline validations rather than model-backed empirical findings. Our work lays the foundation for building more reliable, state-aware RAG agents that resist drift across complex multi-turn interactions.

Read PDF

Similar papers

Preprint Aug 2026

Self-Correcting Long-Horizon Search Agents via Tree-Structured Memory

Experiments establish ReTree as an effective self-correcting memory abstraction for long-horizon search, and show that ReTree consistently outperforms Full-Trajectory ReAct in question-answering and search benchmarks.

Aijun Yang, Qianxue Guo, Ziyi Huang et al. · 0 citations
#large language models Open access Sep 2026

Controlled knowledge updating in memory-enabled AI agents: beyond model editing and recall

Large language model agents that persist across sessions, tools, users, and changing environments do more than answer isolated prompts; they accumulate state. When new evidence arrives, the central question is where that change should live: transient context, external memory, tool or workflow definitions, activation st...

Gabriel Chavira-Juárez, Eder Jahir Gonzalez Bravo, G. Rivera-García et al. · 0 citations
Preprint Aug 2026

HERO: Human-profile Enhanced Retrieval Optimization Framework for Long-term Agent Memory

This work proposes a novel Human-profile Enhanced Retrieval Optimization framework for long-term agent memory (HERO), which converts the dialogue history into a traceable heterogeneous memory graph that preserves raw dialogue text as evidence for reasoning, thereby mitigating information loss.

Yuanhua Lin, Yile Li, Zhiyuan Zhao et al. · 0 citations
Preprint Aug 2026

Navigation-Informed Embeddings: Dense-Retriever Adaptation from Agent Search Traces

Agentic retrieval workflows produce query, retrieval, and stopping traces as a byproduct of answering questions. We study how these traces can adapt a deployed dense retriever to changing workflow distributions without new relevance labels, synthetic queries, or LLM judgments. We introduce Navigation-Informed Embedding...

Shrey B. Shah, Levent Ozgur · 0 citations
Preprint Aug 2026

HyperSkill: Self-Evolving LLM Agents via Hypergraph-Structured Skill Memory

This work proposes HyperSkill, a hypergraph-based memory framework that jointly improves what to store, how memory is structured and retrieved, and how memory evolves, and represents memory as a hypergraph with two node types, subtask steps and reusable skills.

Ruiyao Xu, Tiankai Yang, Wei-Chieh Huang · 5 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.