A scenario-ensemble Carbon Valley dataset for AI-infrastructure cumulative emissions pressure
This repository contains the dataset and reproducibility package for the Data Descriptor “A scenario-ensemble Carbon Valley dataset for AI-infrastructure cumulative emissions pressure.” The dataset supports reproduction, auditing and extension of the Carbon Valley cumulative accounting framework used in the associated study “Rapid artificial intelligence deployment increases near-term pressure on global carbon budgets.” The Carbon Valley is defined as the interval between the onset of artificial-intelligence infrastructure emissions and flow-level breakeven, when effective annual mitigation equals or exceeds annual infrastructure-related emissions. The archive provides the processed inputs, intermediate records, outputs, metadata and executable code required to reproduce the deposited calculations. The archive covers four artificial-intelligence infrastructure deployment pathways — Lift-Off, Base, Headwinds and High Efficiency — over 2024–2035. Each pathway is propagated through 10,000 Monte Carlo scenario realizations, giving 40,000 scenario-level parameter draws. Annual outputs are subsequently evaluated under the full-coupling bounding case, μ = 1.0, and three mitigation-coupling cases, μ = 0.50, μ = 0.25 and μ = 0.10, producing 1,920,000 annual output records. The Monte Carlo layer represents structured scenario uncertainty rather than calibrated probabilities of future outcomes. Uniform distributions are used for specified bounded scenario ranges, while mitigation potential and manufacturing intensity are generated as normal draws and clipped to their prescribed bounds. [+] Clipping leaves a point mass at each bound rather than redistributing the tails, so approximately 16 % of the draws for each of these two parameters take the bound value exactly; the result is not a truncated normal. The same parameter-sampling design is applied across all four deployment pathways; pathway differences arise from their deterministic electricity-demand and installed-capacity trajectories. The accounting framework combines marginality-adjusted operational emissions, stylized embodied-emissions refresh-cycle pulses and logistic mitigation. Annual balances are accumulated by [+] trapezoidal integration, initialised at zero in 2024, as implemented in the deposited code. Two complementary cumulative measures are provided: truncated cumulative positive pressure, which records accumulated gross positive emissions pressure and cannot decline, and a signed cumulative balance, which accounts for both positive and negative annual balances and may decline after flow-level breakeven. These measures answer different analytical questions and should not be interpreted interchangeably. [+] Embodied emissions are represented as stylized refresh-cycle pulses scaled by a fixed boundary-allocation factor, α_emb = 0.00615. This is a calibration and boundary-allocation parameter, not an independently measured fraction of hardware reaching end of life. At this value the embodied term contributes approximately 0.16 % of total infrastructure-related emissions over 2024–2035, so the deposited results are driven almost entirely by the operational term. The archived parameter-sensitivity records confirm this independently: manufacturing intensity and refresh-cycle length have rank correlations with cumulative pressure that are indistinguishable from zero in every pathway. Users interested in embodied carbon specifically should treat this archive as documenting the operational pathway, and should replace the pulse formulation with an explicit additions–retirements–replacement cohort model before drawing embodied-carbon conclusions. The reference-output and calibration records verify that the Lift-Off full-coupling reference path reproduces the published reference target of approximately 2.85 Gt CO₂e of truncated cumulative positive pressure and a flow-level breakeven year of approximately 2031.8. This is a reproduction and calibration check, not an independent validation of the model. The package includes: • harmonized annual scenario inputs and their processed source anchors; • 40,000 Monte Carlo parameter-draw records; • annual infrastructure-emissions, mitigation, balance and positive-pressure outputs; • truncated and signed cumulative trajectories; • reference-output paths for the mitigation-coupling cases; • breakeven, censoring, lag-penalty and other derived indicators; • parameter-sensitivity records based on Spearman rank correlation; • accounting-identity, parameter-bound, completeness, monotonicity, convergence and reproducibility checks; • metadata dictionaries, computational-environment information and file-level manifest; • figures and executable Python code required to regenerate the archived products. [+] Reproduction. The full workflow regenerates the deposited outputs from the archived master seed in approximately 17 seconds, reproducing the annual outputs to a maximum absolute difference of 2.7 × 10⁻¹² across all numeric columns and the published reference path exactly. The interpreter and package versions in which this was verified are pinned in environment.txt; requirements.txt specifies minimum compatible versions only. A file-level SHA-256 manifest is provided in MANIFEST.csv. [+] Appropriate use. The dataset is a stylized global-average analysis. It is not a regional or facility-level inventory, not a life-cycle assessment, and not a forecast, and the electricity module does not resolve data-centre location, dispatch, regional generation mix or changes in marginal generators. Cumulative emissions pressure is an accounting indicator and is not expressed as a share of any remaining carbon budget; users wishing to do so must select and cite a budget, a temperature target and a likelihood level themselves. The full-coupling case μ = 1.0 is a bounding case retained to reproduce the published reference path and should not be reported as a central estimate. Timing statistics should be quoted together with the fraction of realizations reaching breakeven, since censored realizations are neither failures nor zeros. Scenario anchors are derived from processed information from the International Energy Agency, Energy and AI, World Energy Outlook Special Report (2025), with the pathway processing and model parameterization documented in the associated Communications Earth & Environment study and its Supplementary Information. The raw IEA source annex is not redistributed for licensing reasons; all processed scenario values required to reproduce the deposited calculations are provided in the archive and documented in the accompanying Data Descriptor. This revised release was prepared in response to peer review. It improves provenance documentation, aligns the mathematical description with the deposited implementation, adds the signed cumulative measure and supplementary sensitivity records, documents the computational environment, [+] resolves the divide-by-zero warnings reported during review (the guard changes no model output; patched and unpatched runs are bit-identical), and strengthens reproducibility checks. The underlying four-pathway scenario architecture, Monte Carlo design, master random seed (24042026), and published reference calibration are unchanged.