Difference-in-differences (DID) studies are widely used to evaluate climate policy, but assessing the evidence supporting their identification assumptions remains challenging. We introduce ARGUS, a structured language-model pipeline that audits reported evidence against an eleven-dimension assumption-implication-evidence rubric and abstains when relevant evidence cannot be retrieved. We evaluate ARGUS using injected flaws, economics papers, and a small pilot with reconciled labels. On the 11-flaw benchmark, ARGUS detects 73% of planted flaws, compared with 18% for a keyword-based pipeline. Across 26 economics papers, ARGUS abstains on about 40% of paper-dimension assessments for lack of retrievable evidence. In a five-paper pilot with labels reconciled by two annotators, it assigns a higher risk level than the labels on 25 of the 33 assessments it completes. A rule fixed before the labels arrived removes most of this in-sample; weighted agreement stays low. ARGUS provides evidence-linked risk reports that localize potential weaknesses for expert review, without adjudicating causal claims. Code and data: https://github.com/yonghongzhang-io/ARGUS
Abstract Official statistics have long provided a foundation for evidence-informed policy-making, offering professionally independent, quality-assured and transparent evidence to support democratic decision-making. Yet the contemporary policy environment presents statistical systems with a difficult set of pressures. D...
Steve Macfeely, Ashley Ward· Administration· 1 citation
Causal inference plays a central role in sustainability studies, providing the foundation for evidence-based scientific discovery and policy evaluation. However, uncovering credible causal relationships in complex real-world settings often relies heavily on the judgment of domain experts and extensive contextual knowle...
Mengying Zhang, Chao Fan· Environmental Research Commu...· 0 citations
V VERA-8B is a new end-to-end audit reasoning system that identifies audit risks before enforcement actions occur, and is the first to unify SFT and GRPO for evidence-grounded audit reasoning under one evidence standard.
GPS-Bench is introduced, an evidence-grounded benchmark for governance policy simulation that links policies to relevant actors, actor actions and downstream impacts using legislative records, lobbying disclosures, regulatory documents, corporate filings, economic data and other public evidence.
Linh Le, Melanie Bui, My Chiffon Nguyen et al.· 0 citations
SAFE, a controlled benchmark in which models make deployment decisions with optional evidence that varies in retrieval cost, probability, severity, and presentation, is introduced and suggests that deployment-time safety depends not only on how models respond to known risks, but also on whether they acquire the evidenc...
Deep Research agents synthesize evidence into cited reports, yet a well-cited report can still reach a misleading conclusion. Citation correctness checks whether cited sources support individual claims. It does not show whether adaptive search exposed a representative view of all documents made available for evaluation...
Shu-Yao Xiao, Sheng-Ling Wang, Xuan Chen et al.· 0 citations
Related blog posts
MIT News · Artificial Intelligence· news.mit.eduSep 24, 2026
With millions of users across the world, Julia has been used to conduct cutting-edge research and to design new drugs, jet engines, heat pumps, and more.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.