Skip to content

CausalArena: Benchmarking Causal Discovery in the Foundation Model Era

Sep 2026 · 0 citations · 46 references
Computer Science

TL;DR

CausalArena is introduced, a unified and evolvable benchmark for causal discovery under a common protocol, and substantial ranking shifts across SCM families and protocols are revealed, showing that strong performance in one benchmark regime does not reliably transfer to others.

Abstract

Causal discovery aims to uncover causal structures from data and is fundamental to scientific reasoning and intervention-based decision making. Its evaluation relies heavily on structural causal models (SCMs), which specify a causal graph together with the mechanisms that generate data, yet existing studies differ substantially in graph families, mechanisms, and evaluation protocols. The emergence of causal discovery foundation models (CDFMs) further complicates evaluation: performance may reflect not only causal discovery ability, but also overlap between pretraining environments and test SCMs, making results on fixed synthetic benchmarks difficult to interpret. We introduce CausalArena, a unified and evolvable benchmark for causal discovery under a common protocol. Synthetic SCMs supply controlled breadth over structures and mechanisms; semantic operational SCMs provide human-auditable, semantically grounded environments beyond standard synthetic generators; and formula-grounded SCMs test discovery under explicit scientific mechanisms. Public real-world datasets provide an additional external-validity check. Experiments across classical, neural, and pretrained methods reveal substantial ranking shifts across SCM families and protocols, showing that strong performance in one benchmark regime does not reliably transfer to others. These results highlight benchmark diversity and pretraining--evaluation overlap as central challenges for evaluating causal discovery in the foundation model era.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

CausalVerify: An Execution-Grounded Benchmark for LLM Causal Inference Workflows

Existing causal-inference benchmarks for LLMs mostly score method descriptions or whether generated code runs, not whether the executed workflow recovers the target causal estimate. CausalVerify studies this verification problem for structured econometric causal-estimation workflows by separating realistic interpretati...

Yong-Hong Zhang, Ricardo Correia, Isabel M. Parra et al. · 1 citation
Preprint Aug 2026

Interpretable Causal Discovery via Causal-Effect Constraints

This work considers the task of conditional causal discovery as a Bayesian inference problem, in which the posterior is targeted over causal graphs and parameters conditional on an event such as a causal-effect constraint, and adapts rare-event estimation techniques to perform inference the joint graph-parameter space.

Cixuan Zhang, Guy Van den Broeck, Benjie Wang · 0 citations
Conference Open access Sep 2026

THGAgents: Traceable Biomedical Hypothesis Generation via Dynamic Causal Reasoning

THGAgents utilizes collaborative and dynamically updating agents to build a Traceable Causal Knowledge Graph, which serves as the foundation for the evidence-based knowledge structure and employs an LLM-driven heuristic search algorithm to traverse the complex network, balancing both novelty and rigor to deduce strict,...

Ming-Jia Yang, Kun-Hua Dong, K. Lim et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Can LLMs Discover Scientific Laws in Real and Parallel Worlds?

It is found that predictive fit can diverge from scientific validity, memorization shapes whether models reproduce or move beyond published formulas, and the best-of-N study reveals a selection bottleneck.

Yi-Ming Huang, Zi-Chen Liu, Junxia Cui et al. · 1 citation
#software testing Review Sep 2026

Statistical Inference for Bivariate Functional Causal Discovery

This paper formalizes a test-based approach for bivariate causal discovery by repurposing goodness-of-fit and independence tests within a hypothesis-testing framework and demonstrates the use and behavior of the inferential framework through simulations that vary the degree of assumption violation, as well as through r...

Shreya Prakash, Fan Xia, Elena Erosheva · 0 citations
#artificial intelligence Preprint Sep 2026

Active Causal Discovery Benchmark: Evaluating LLM Agents Under Budgeted Interventions

We introduce the Active Causal Discovery Benchmark (ACDB), an SCM-grounded environment for evaluating whether LLM agents recover causal graph structure from observations and budget-constrained hard interventions. ACDB pairs a linear-Gaussian world generator with a fixed observe-intervene-submit API and a three-layer sc...

Sagar Deb, Devam Shah, Ashwanth Krishnan · 0 citations

Related blog posts

MIT News · Artificial Intelligence Oct 7, 2026

Discovering the value of humanistic inquiry

Students in MIT’s Concourse program delve deeply into the human condition, debate challenging questions, and learn to develop judgment about issues that can’t be quantified.

Microsoft Research Blog Oct 7, 2026

Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses

Training AI agents with reinforcement learning can be challenging because their tools, context, and decision-making are managed by complex frameworks. Agent Lightning connects existing agents to RL training, making it easier to improve them without rebuilding them. The post Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.