(Invited) Autonomous, Data-Driven Acceleration of Energy Storage Technology, from Electrons to Devices
Electrochemical energy conversion and storage sit at a decisive inflection point: urgency of breakthrough is high, yet the science remains dominated by multiscale physics, interfacial complexity, sparse high-fidelity data, and long validation cycles for safety and lifetime. I have a pragmatic thesis: AI does not “replace” electrochemistry; it accelerates it when models are anchored to mechanistic constraints, uncertainty is taken care of, and workflows are designed to close loops between simulation, experiment, and optimization. I will discuss a disciplined operating model for scientific machine learning in batteries: defining the decision that must be improved (screening and discovery), choosing representations that encode physics, and quantify where the model can and cannot be trusted under distribution shift. We want to compress the “electrons-to-devices” pipeline by replacing repeated expensive computations and measurements with calibrated surrogates, while using active learning and multi-fidelity strategies to allocate scarce high-quality labels where they matter most. A core effort addresses electronic-structure bottlenecks that propagate into materials screening, interfacial chemistry, and kinetics. We introduce equivariant graph neural network frameworks that predict electron density across molecules, liquids, and solids at speeds orders of magnitude beyond density functional theory (DFT). To reduce data load and improve practicality, we shift the learning target from dense real-space grids to sparse, SE(3)-equivariant density matrices, reaching high accuracy with far smaller datasets while enabling uncertainty diagnostics tied to charge error and self-consistency. We have extended compact electronic representations further by learning “floating orbital” placements and coefficients through an equivariant Cartesian tensor network with controlled symmetry breaking; beyond accuracy, these models deliver immediate computational value by reducing self-consistent-field iterations when used to initialize DFT. Scaling upward, we develop and operationalize uncertainty-aware machine-learned interatomic potentials (MLIPs), including general-purpose foundation models for stable molecular dynamics across diverse chemistries, alongside evidence-based limits of universality in energy materials. A case study in polyanion sodium cathodes demonstrates that out-of-distribution performance often requires system-specific datasets and charge-aware architectures; fine-tuning universal MLIPs can recover much of the deficit, but domain-specialized models can remain faster and more accurate. To make specialization economical, we have done autonomous batch active learning, integrating uncertainty, batch selection, and efficient force/stress evaluation, to reduce labeling cost and expert intervention. We also advance data efficiency at the quantum-chemistry level through minimal multilevel machine learning which optimally allocates training across fidelity hierarchies to minimize combined prediction error and acquisition wall-time. Finally, to break the equilibrium bias that has limited reactive modeling, we curate pathway-rich datasets and show that equivariant models trained on such data can serve as practical surrogate potentials for reaction-path search, enabling kinetics at scales previously infeasible; these concepts are demonstrated in electrochemical interface kinetics and in long-timescale cathode dynamics where nanosecond molecular dynamics becomes accessible while retaining DFT-level fidelity. Our effort moves from forward prediction to inverse design and autonomy, where the true payoff is not a faster model, but a faster, safer, more reliable decision. For battery interphases such as the solid-electrolyte interphase (SEI), we outline a blueprint for inverse design that couples multi-scale, multi-fidelity simulations with operando characterization, high-throughput synthesis/testing, and semi-supervised generative learning. The essential ingredient is uncertainty tracking across all layers, experimental, computational, and generative, so fidelity can be improved deliberately, and “failed” experiments can be used to reduce bias and expand coverage. More broadly, we develop and deploy generative methods for molecules and materials, diffusion models for organometallic catalysts and crystalline structures, symmetry-consistent generation for disordered crystals, and data-free or amortized global optimization—paired with hierarchical screening (“materials funnels”) that combine generative proposal, fast learned surrogates, MLIPs, and final DFT validation. Because acceleration without trust is fragile, we place calibrated uncertainty at the center: ensemble-based probabilistic predictors unify aleatoric and epistemic components and are recalibrated on unseen data; early-aging models forecast full degradation trajectories from limited battery cycling with uncertainty and explainability sufficient to support uncertainty-guided truncation of long tests; and sensitivity analysis of physics-based degradation models via differentiable surrogates identifies the parameters that truly govern capacity fade. We take an integrated view of distributed, intention-agnostic materials acceleration platforms that orchestrate asynchronous experiments and simulations across institutions. The overarching message is operational: when physically grounded representations, calibrated uncertainty, data-efficient multi-fidelity learning, and closed-loop autonomy are engineered as a single system, electrochemical science can reduce time-to-insight and time-to-validation, maintaining interpretability and reproducibility.