Language models increasingly perform empirical analyses end to end, yet existing evaluations assess the written explanation or whether generated code executes, not whether the executed workflow recovers the intended causal estimand. We introduce CausalVerify, an execution-grounded benchmark for end-to-end causal analys...
Yong-Hong Zhang, Ricardo Correia, Isabel M. Parra et al.· 0 citations
Difference-in-differences (DID) studies are widely used to evaluate climate policy, but assessing the evidence supporting their identification assumptions remains challenging. We introduce ARGUS, a structured language-model pipeline that audits reported evidence against an eleven-dimension assumption-implication-eviden...
Yong-Hong Zhang, Yong Xie, Isabel M. Parra et al.· 0 citations
The results suggest that benchmark scores should be interpreted together with their scaffolding level, scoring criterion, and reliability profile, providing a practical framework for more valid evaluation of LLM agents.
Yong-Hong Zhang, Shadi Motaali, Vu Phong Dinh et al.· 2 citations
Existing causal-inference benchmarks for LLMs mostly score method descriptions or whether generated code runs, not whether the executed workflow recovers the target causal estimate. CausalVerify studies this verification problem for structured econometric causal-estimation workflows by separating realistic interpretati...
Yong-Hong Zhang, Ricardo Correia, Isabel M. Parra et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.