Skip to content

GANDR: Claim Auditing for Verifiable Legal Answer Generation

Sep 2026 · 0 citations · 35 references
Computer Science

TL;DR

GANDR (Grounded ANswer DRafter), a two-agent system in which a Drafter writes an answer in a structured legal-reasoning format and a separate Critic, with the same view as a human verifier, audits each claim against its cited source and emits a per-claim audit trace on every round.

Abstract

In high-stakes domains such as legal practice, a language-model answer is only useful to the extent that a reader can verify each claim against the source the system cites. Current grounded-generation pipelines score the answer as a whole, so a correct conclusion can rest on fabricated or loosely matched citations and still score well. Closing this gap requires both a system built for per-claim verification and an evaluation that measures it. We introduce GANDR (Grounded ANswer DRafter), a two-agent system in which a Drafter writes an answer in a structured legal-reasoning format and a separate Critic, with the same view as a human verifier, audits each claim against its cited source and emits a per-claim audit trace on every round. We pair it with a strict correctness criterion requiring every citation to resolve to a passage the retriever returned. On a 185-item legal benchmark where all six systems share one backbone, one retrieval surface, and one citation instruction, GANDR ranks first on every primary metric, reaching 70.8% strict accuracy and leading the strongest baseline by 11.3 points (p<0.01). Reverting the protocol-anchored commit rule lowers strict accuracy by 22.7 points, and the strict lead stays positive on three further backbones, at +3.2 to +6.5 points. This lead traces to the Drafter configuration and the protocol-anchored commit, not to rewriting. Against two law-trained annotators the audit flags under-supported claims at F1 0.84 as a binary detector, while its four-way verdict labels agree only weakly and are advisory. Code is available upon request.

View source

Similar papers

Preprint Aug 2026

TRACES: A Benchmark for Epistemic Reliability in Scientific Reasoning by LLMs

A probe corpus of 42 retracted, fraudulent, and pseudoscientific papers is paired with a methodology for eliciting and scoring single-shot model engagement with each paper's framing, indicating an urgent need for guardrail infrastructure for scientific deployment of language models.

V. Rodionov, Shamil Assylbekov · 0 citations
#artificial intelligence Review Sep 2026

Claim-Gated Source-Risk Auditing for Generative Search

A generative search answer can cite a supported passage yet omit a source relationship that changes its interpretation. We specify a claim-gated audit of the query-source-answer tuple. An omission is resolved only when relationship evidence, answer adoption, materiality, and disclosure are all observed; incomplete evid...

Kai-Nan Zhou, Chu-Hong Xu, Gang-Zhen Qian et al. · 0 citations
Open access Sep 2026

EviGuard: Machine-Verifiable Evidence Grounding for LLM-Based Industrial Incident Reasoning

EviGuard is presented, a system that decides when an LLM’s understanding is trustworthy enough to act on and has an ensemble of deterministic verifiers label every claim supported, contradicted, or unknown against the graph—honoring interval time, event-time policy and credential versions, network reachability, and phy...

Hao-Zhe Zhou, Hang Lei, Mao-Lin Yang · 0 citations
#artificial intelligence Preprint Sep 2026

Legal Research Bench: Measuring End-to-End Reliability in Long-Horizon Legal Research Agents

Legal research is a core and time-consuming legal workflow. Lawyers must identify controlling authority, verify that it remains valid, reconcile statutes and cases, and synthesize a grounded answer. Language model agents are a natural fit for this retrieval-intensive workflow, and automating even part of it would be va...

К. А. Дроздов, Oliver Chen, Langston Nashold et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Ingest-Time Fact Compilation for Cost-Efficient and Reliable Question Answering over Revised Corpora

This work presents ingest-time fact compilation, an architecture that performs this work when corpus data is ingested or changed and supports a narrow but practical claim: resolving a corpus state once can make subsequent QA cheaper and more reliable for inexpensive models.

Kyle Wild, Yusuke Takahashi, Asako Uraki · 0 citations
#artificial intelligence Preprint Sep 2026

What the Window Does Not Contain: Auditing Provenance in a Document-Grounded Instability Benchmark

Ask a language model the same question about the same document twenty times, and it sometimes returns two different answers. We built Probity, a benchmark of 60 tasks and 470 items from real venture-financing filings, to measure how often this happens. Then we audited our own corpus and found a defect any excerpt-built...

Seyed Mosayeb Alam · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.