Skip to content
Review Open access

Evaluating context in LLM prompts for causal inference in empirical sustainability studies

Aug 2026 · Environmental Research Communications · Vol 8 · 0 citations · 41 references
Physics

Abstract

Causal inference plays a central role in sustainability studies, providing the foundation for evidence-based scientific discovery and policy evaluation. However, uncovering credible causal relationships in complex real-world settings often relies heavily on the judgment of domain experts and extensive contextual knowledge. With the emergence of large language models (LLMs), an open question arises as to whether these models can meaningfully support causal reasoning for sustainability research. This study systematically evaluates the capacity of LLM to infer causal hypotheses and methodological structures from contextual descriptions of empirical research. We construct a benchmark based on five high-quality, peer-reviewed sustainability studies and design structured prompts that replicate the informational context available to human researchers prior to estimation. LLM outputs are evaluated against expert-validated ground truth in terms of causal edge recovery, scope expansion, and alignment with identification strategies, allowing for a quantitative assessment of conceptual causal reasoning rather than numerical estimation. Our findings indicate that GPT-5 recovers core causal edges with moderate accuracy (0.33–0.85), and assigns causal direction with high reliability (0.87–1.00). However, the scope expansion rate achieves roughly 0.58–0.93 of the causal edges proposed by the model, indicating a strong tendency toward over-connection between variables. Overall, our findings contribute to a deeper understanding of the potential and limitations of LLMs as tools for causal reasoning and methodological support in empirical sustainability research.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.