Skip to content

Ingest-Time Fact Compilation for Cost-Efficient and Reliable Question Answering over Revised Corpora

Aug 2026 · 0 citations · 17 references
Computer Science

TL;DR

This work presents ingest-time fact compilation, an architecture that performs this work when corpus data is ingested or changed and supports a narrow but practical claim: resolving a corpus state once can make subsequent QA cheaper and more reliable for inexpensive models.

Abstract

Most agentic question answering (QA) systems do an important part of their semantic work at the worst possible time: every time someone asks a question. When a corpus contains revisions, drafts, revocations, deletions, and sources with different levels of authority, the model must reconstruct the governed current state on every read - then throw that work away and repeat it on the next query. This is a bit like a database that rebuilds a materialized view every time someone reads from it. We present ingest-time fact compilation, an architecture that performs this work when corpus data is ingested or changed. Raw passages are rephrased into self-contained facts; rules governing revisions, deletions, effective dates, and source trust are resolved once; and the resulting state is stored as typed records carrying source and revision provenance. At query time, an inexpensive model reads the compiled record instead of reconstructing it from noisy candidates. In a controlled synthetic experiment across five seeds, the same low-cost model produced the correct value, source, and revision in only one of 30 trials under query-time reconstruction, but in all 30 trials from the compiled substrate, at 12.89 times lower mean read cost per question. On simpler revision questions both architectures were exact, but the compiled path used 21.6 times fewer tokens. A separate test found that fact rephrasing roughly halved verbose Federal Reserve dialogue while preserving high source entailment, but left concise Wikipedia prose essentially unchanged. These results support a narrow but practical claim: resolving a corpus state once can make subsequent QA cheaper and more reliable for inexpensive models. We release the open source, MIT-licensed implementation and experimental artifacts.

View source

Similar papers

Preprint Aug 2026

RAG Deserves an Index: Why Ingest-Time Compilation Beats Query-Time Interpretation

The paradigm ingest-time semantic compilation (ISC) is called: compile a corpus's meaning into a queryable substrate with two coupled layers - incrementally maintained embeddings, and atomic claims whose provenance is validated at compile time - and treat that substrate as a first-class database object with its own DDL...

Kyle Wild, Yusuke Takahashi, Asako Uraki · 0 citations
#artificial intelligence Preprint Sep 2026

Version- and Scope-Aware Question Answering over Normative Documents: A Deployed System and an End-to-End Evaluation at Production Scale

Correctly answering a question grounded in normative documents often depends on information outside any single passage: whether the retrieved document is the version currently in force; whether it applies to the jurisdiction, subject (such as an institution or applicant), and date at issue; and whether each normative c...

Liuyin Wang, Shuai-Peng Jin, Ji-Wei Shi et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Contract Memory Compiler: Resolve, Then Traverse

External memory lets language-model agents answer questions about histories too long for the answer model's context window. Updates create a harder problem than retrieving a recent fact: changing one relation can redirect a multi-hop question to records about an entity absent from the question. We study this update-dep...

Zhi Song, Ximing Xing, Chun-Han Li et al. · 2 citations
Preprint Aug 2026

Time Present and Time Past: Benchmarking Large Language Models on Temporally Evolving Document Understanding

Evolving documents, such as laws, tax codes, and software documentation, are amended, replaced, and sometimes reverted over time, so a question has different correct answers at different dates. In contrast to encyclopedic knowledge, where an old fact is simply overwritten, an amendment is itself an official text that s...

M. Sobhani, Md. Faiyaz Abdullah Sayeedi, F. Chowdhury et al. · 1 citation · ⚡1
Preprint Aug 2026

TraceSQL: Traceable Answerability Estimation for Reference-Free Text-to-SQL Verification

Results show that SQL verification can be performed with a lightweight learned model while retaining feature-level evidence for inspecting and diagnosing its predictions, and feature attribution shows that the model relies on both semantic grounding and deterministic SQL-structure signals.

N. Shukla, Debasmita Panda, Srutanik Bhaduri et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.