Skip to content

Evidence-Grounded Auditing of Identification Assumptions in Climate-Policy Causal Evaluations

Sep 2026 · 0 citations · 52 references
Computer Science

Abstract

Difference-in-differences (DID) studies are widely used to evaluate climate policy, but assessing the evidence supporting their identification assumptions remains challenging. We introduce ARGUS, a structured language-model pipeline that audits reported evidence against an eleven-dimension assumption-implication-evidence rubric and abstains when relevant evidence cannot be retrieved. We evaluate ARGUS using injected flaws, economics papers, and a small pilot with reconciled labels. On the 11-flaw benchmark, ARGUS detects 73% of planted flaws, compared with 18% for a keyword-based pipeline. Across 26 economics papers, ARGUS abstains on about 40% of paper-dimension assessments for lack of retrievable evidence. In a five-paper pilot with labels reconciled by two annotators, it assigns a higher risk level than the labels on 25 of the 33 assessments it completes. A rule fixed before the labels arrived removes most of this in-sample; weighted agreement stays low. ARGUS provides evidence-linked risk reports that localize potential weaknesses for expert review, without adjudicating causal claims. Code and data: https://github.com/yonghongzhang-io/ARGUS

View source

Similar papers

Open access Sep 2026

Evidence-informed policy-making in a data-driven age: The role and limits of official statistics

Abstract Official statistics have long provided a foundation for evidence-informed policy-making, offering professionally independent, quality-assured and transparent evidence to support democratic decision-making. Yet the contemporary policy environment presents statistical systems with a difficult set of pressures. D...

Steve Macfeely, Ashley Ward · 1 citation
Review Open access Aug 2026

Evaluating context in LLM prompts for causal inference in empirical sustainability studies

Causal inference plays a central role in sustainability studies, providing the foundation for evidence-based scientific discovery and policy evaluation. However, uncovering credible causal relationships in complex real-world settings often relies heavily on the judgment of domain experts and extensive contextual knowle...

Mengying Zhang, Chao Fan · 0 citations
#artificial intelligence Preprint Sep 2026

GPS-Bench: A Governance Policy Benchmark for Automating Policy Analysis

GPS-Bench is introduced, an evidence-grounded benchmark for governance policy simulation that links policies to relevant actors, actor actions and downstream impacts using legislative records, lobbying disclosures, regulatory documents, corporate filings, economic data and other public evidence.

Linh Le, Melanie Bui, My Chiffon Nguyen et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Do Frontier Models Seek Safety Evidence Before Acting?

SAFE, a controlled benchmark in which models make deployment decisions with optional evidence that varies in retrieval cost, probability, severity, and presentation, is introduced and suggests that deployment-time safety depends not only on how models respond to known risks, but also on whether they acquire the evidenc...

Omer Tafveez · 0 citations
#artificial intelligence Review Sep 2026

Search Shapes Conclusions: Auditing Evidence Selection Bias in Deep Research Agents

Deep Research agents synthesize evidence into cited reports, yet a well-cited report can still reach a misleading conclusion. Citation correctness checks whether cited sources support individual claims. It does not show whether adaptive search exposed a representative view of all documents made available for evaluation...

Shu-Yao Xiao, Sheng-Ling Wang, Xuan Chen et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.