TimelyRAG is proposed, a retriever-agnostic framework that incorporates temporal distance into ranking to align queries with version-appropriate documents, and TimelyQABench is introduced, the first benchmark for regulation-heavy domains with overlapping-evolving challenges.
Abstract
Although large language models (LLMs) and retrieval-augmented generation (RAG) have advanced open-domain question answering (QA), they remain unreliable when documents evolve through amendments. Existing time-sensitive retrieval methods address only the disjoint-evolving environment, where each update is an independent snapshot. However, laws, policies, and regulations often operate in overlapping-evolving environments, where amendments override earlier clauses while preserving most content, creating strong semantic overlap across versions. We propose TimelyRAG, a retriever-agnostic framework that incorporates temporal distance into ranking to align queries with version-appropriate documents. We also introduce TimelyQABench, the first benchmark for regulation-heavy domains with overlapping-evolving challenges. Experiments show consistent gains, up to +28.6% in nDCG@10, highlighting the importance of temporal reasoning for reliable QA over evolving documents. All resources are available at https://github.com/kaist-dmlab/TimelyRAG.
This work introduces TempFinRAG, a point-in-time evaluation protocol built from public filings and XBRL facts, and introduces TempFinQA, a point-in-time evaluation protocol built from public filings and XBRL facts, and evaluates the framework on complementary evidence-grounded, numerical, conversational, and multi-tabl...
Lanju Tao, Zheng-Ji Li, Ying-Rui Ji et al.· Symmetry· 0 citations
Temporal Information Retrieval (TIR) has been increasingly critical given the rise of Retrieval-Augmented Generation (RAG). Since temporally mismatched evidence can be highly misleading, TIR aims to retrieve documents that are both semantically and temporally relevant to a query. Two TIR paradigms have emerged - tempor...
Soyeon Kim, Hyunjin Kim, J. Bak et al.· 0 citations
Results align with a diagnostic perspective on chunking: using evidence at a task-appropriate level of granularity can improve grounding, auditability, and answer quality, but the observed patterns should be interpreted within the HotpotQA distractor setting, fixed generator, and tested context budgets.
Large Language Models (LLMs) offer strong capabilities for Natural Language Processing, yet their inherent uncertainty often produces hallucinations, confident but incorrect statements, which is critical in domains requiring precise knowledge representation. Retrieval-Augmented Generation (RAG) reduces this risk throug...
Alexandra V. Jove-Ticona, Luis J. Duarte-Coaquera, Israel N. Chaparro-Cruz et al.· International Journal of Adv...· 0 citations
W-RAG is proposed, a source-aware retrieval framework that performs ontology-guided retrieval, local ranking within each knowledge base, and source-level weighting to regulate evidence composition to improve document coverage and generation quality.
Hridya Dhulipala, Rajesh Ombase, Michael Wang et al.· 0 citations
This work proposes EnSI-RAG (Entity-Structure-Indexed Retrieval-Augmented Generation), a framework that constructs a query-independent, entity-centered index that separates evidence localization from answer synthesis while preserving traceable source evidence.
Xuan-Yu Meng, Jiashuo Sun, Jash Parekh et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.