An integrated workflow that combines document-grounded large language model (LLM) literature mining, human verification, and probabilistic modeling to enable uncertainty-aware prediction of the coefficient of thermal expansion (CTE) in complex oxides is introduced.
Abstract
Machine learning is increasingly used in materials discovery, but its practical application is often limited by the time required to construct structured experimental datasets and by the lack of reliable uncertainty estimates. We introduce an integrated workflow that combines document-grounded large language model (LLM) literature mining, human verification, and probabilistic modeling to enable uncertainty-aware prediction of the coefficient of thermal expansion (CTE) in complex oxides. Experimentally reported CTE values, compositions, and measurement temperature ranges are extracted from full-text articles using a document-grounded large language model and subsequently normalized and verified. Composition-derived descriptors and the reported temperature bounds are used as inputs to a probabilistic regression model that predicts both the expected CTE and a composition-dependent uncertainty, enabling prediction intervals for new compositions. On held-out tests, the model achieves competitive mean accuracy relative to deterministic baselines while producing uncertainty estimates that increase systematically for sparsely represented or chemically distinct compositions, enabling risk-aware screening and prioritization. The workflow supports comparative screening of compositions with targeted CTE behavior and helps guide experimental selection for detailed thermophysical characterization. This study illustrates how LLM-assisted literature curation can be combined with uncertainty-aware machine learning to construct property prediction workflows for materials systems where experimental data are sparse and primarily available in the literature.
Machine-learned interatomic potentials (MLIPs) have become the state-of-the-art for performing accurate, scalable molecular dynamics (MD) simulations. It is, therefore, crucial to understand and quantify the reliability of MLIPs for downstream property predictions. Uncertainty in predicted properties can arise from limitations in first-principles training data, intrinsic MLIP model errors in representing the data, and the statistical noise introduced during subsequent MD simulations. Using ion transport in Li7P3S11 as a case study, we systematically assess the impact of training set size and selection, neural network stochasticity, and MD sampling statistics on predicted diffusivity and activation energy. We find that when using equivariant MLIP architectures with standard MD protocols, uncertainty arising from MD sampling dominates over model-induced errors. In contrast, MLIP errors relative to the underlying first-principles data are consistently minor. Given this, there are two main routes to improving the accuracy of predictions based on MLIP potentials: adopting higher accuracy reference data generation methods and improving the MD sampling statistics.
Tawfiqur Rakib, Lucas K. Wagner, Elif Ertekin· APL Machine Learning· 0 citations
Accurate prediction of CO2 solubility in formation brines is central to carbon storage design because dissolution trapping reduces CO2 mobility and supports long term containment. Yet, solubility data and correlations are often limited in coverage, uncertain at high salinity and pressure, and can be unreliable when extrapolated beyond the calibration range. This work develops a physics-informed benchmarking framework that evaluates when machine learning (ML) models provide reliable CO2 solubility predictions under reservoir relevant conditions, and when established physics-based correlations remain the safer choice. A physics-informed CO2 brine dataset was generated over geologically realistic ranges that represent deep saline reservoirs at approximately 4,900 to 13,000 ft depth, spanning 35 to 347 bar, 285 to 430 K, and 0 to 259 g/L salinity. Ground-truth solubilities were produced using Henry's law with van't Hoff temperature dependence and a Setchenov salting-out correction, then supplemented with fugacity and activity-coefficient adjustments to address non-ideal behavior at higher pressure and salinity. A calibrated non-ideality term was included to preserve physically consistent monotonic trends across the full Pressure-Temperature-Salinity(P–T–S) space. To emulate laboratory uncertainty without changing the sampled inputs, controlled zero-mean Gaussian noise of 1%, 3%, 5%, and 10% relative standard deviation was applied to solubility targets. Three ML models, Linear Regression, Random Forest, and a three-layer MLP (64-32-16, ReLU), were trained using identical feature sets (P, T, S), standardized preprocessing, and consistent train, validation, and test splits. Model performance was evaluated using R2, MAE, RMSE, k-fold cross-validation, parity and residual diagnostics, calibration curves, and bootstrap uncertainty estimates. Out-of-distribution (OOD) robustness was quantified by withholding a high-salinity band (S > 200 g/L) and a low-temperature band (T < 300 K) from training to test generalization in regimes relevant to storage screening. The resulting workflow provides a practical, physics-consistent basis for selecting solubility predictors and defining reliable application envelopes for ML in CCUS studies.
O. Ejehu, A. J. Whitcomb, M. Hunter et al.· SPE Nigeria Annual Internati...· 0 citations
An explainable machine-learning framework was developed for dielectric constant prediction using 52,168 crystalline materials extracted from the Joint Automated Repository for Various Integrated Simulations (JARVIS-DFT) database, demonstrating the complementary roles of electronic structure and elemental chemistry.
D. Pundhir, Ashok Kumar· Applied Physics A· 0 citations
Machine learning methods for predicting the electron ionization mass spectra from molecular structures have shown promise for environmental chemical identification, but their performance under domain-specific data scarcity remains poorly understood. We systematically compare a conventional multilayer perceptron model (NEIMS) with a Transformer-based chemical foundation model (MolFormer-XL) for the electron ionization mass spectrometry spectrum prediction under controlled few-shot conditions. Using fluorine-containing molecules as a broader proxy domain, including a PFAS-like subset, motivated by the practical challenge of detecting novel fluorinated contaminants with limited reference data, we vary the number of domain-specific training examples from 5 to 175 while maintaining fixed validation and test sets. Across all few-shot conditions and three of four evaluation metrics (weighted cosine similarity, intensity-weighted precision, and top-10 precision), MolFormer-XL consistently outperforms NEIMS, while intensity-weighted recall remains comparable between the two models. The largest performance gaps are observed in extreme data-scarcity regimes. These results demonstrate that MolFormer-XL, which combines pre-trained molecular representations with a Transformer-based architecture and learned SMILES embeddings, provides a promising approach for transfer under severe domain-specific data scarcity in environmental mass spectrometry.
Kinetic model discovery is a central challenge in chemical engineering, as accurate rate expressions are essential for understanding and controlling chemical and biological processes. Symbolic regression (SR) has emerged as a powerful data-driven approach for identifying interpretable kinetic models, but usually operates without domain knowledge, often exploring physicochemically implausible models. Large language models (LLMs) offer a promising avenue for injecting domain expertise into this search. Here, we introduce an LLM-guided SR framework, embedding an LLM module within an iterative SR algorithm for automated kinetic model discovery. The LLM performs two roles at each iteration: (1) a qualitative physicochemical critique of the best SR candidates, and (2) the proposal of new candidate rate expressions guided by the SR-generated models and embedded chemical knowledge. Our framework is evaluated on four in silico case studies of increasing complexity, spanning heterogeneous catalysis and bioprocess systems. Results show the LLM-guided framework reduces iterations to identify the ground-truth model by $41.7-79.3\%$ versus a state-of-the-art SR framework, with the LLM directly proposing the correct model structure in over half of the guided runs. In practical settings, where each iteration typically requires a new wet-lab experiment, this translates into a substantial reduction in experimental effort. Predictive performance on an independent validation set is equivalent between both approaches, with $R^2>0.98$ in all case studies. Ablation studies indicate that both the SR component and the LLM scale contribute to this performance, with a reduced-size LLM largely retaining discovery efficiency. These findings demonstrate that LLMs can effectively inject domain knowledge into scientific model discovery, paving the way toward fully automated, domain-aware kinetic modelling pipelines.
Roberto Aliaga Medina, Paulina Quintanilla, Antonio E. del-Rio Chanona· 0 citations
Materials prediction depends critically on how scientific knowledge is represented, yet many governing considerations exist only as natural-language heuristics that conventional learners cannot use. We introduce CRISP, a large language model-assisted framework that treats representation construction as a rule-space exploration and compilation problem: it repeatedly samples target-relevant chemical rules without access to structures, labels or data splits, consolidates related concepts, and compiles each into an executable scalar descriptor supplied to a conventional learner. For positive-unlabeled inorganic-crystal synthesizability, CRISP outperformed expert-curated and generic structural representations under a shared learner and surpassed purpose-built synthesizability models, with its advantage most pronounced under structural-size and chemical-family shifts. Infrequently generated rules contributed complementary predictive information, showing that generation frequency does not determine utility. The same workflow yielded competitive representations for formation energy and ionic conductivity while revealing task-dependent limits for shear modulus, establishing a dataset-blind, auditable route from broad chemical knowledge to transferable computational representations.
Jaehwan Choi, Kunik Jang, Seongmin Kim et al.· 0 citations