Skip to content
Open access

Closed-Loop Solid-State Synthesis Planning for Materials Discovery With Large Language Models.

Aug 2026 · Advances in Materials · pp. e74502 · 0 citations · 29 references
Medicine

Abstract

Developing reliable synthesis routes for complex materials remains a major bottleneck in accelerating materials discovery. This study establishes a large language model-based framework for predicting and optimizing synthesis conditions directly from the literature data. Key synthesis information, including target compounds, precursors, and processing parameters, was systematically extracted from 4407 open-access solid-state synthesis papers and organized into a structured recipe dataset. Using a retrieval-augmented generation (RAG) approach, the system first retrieves similar recipes from the corpus and then generates a new candidate recipe conditioned on those exemplars. The generated recipes were benchmarked against literature data using quantitative scoring metrics, achieving strong agreement with experimentally reported conditions. To validate the predictive capability, the framework was applied to unreported solid-state electrolyte candidates identified through first-principles screening, and multiple oxy-selenide compounds were successfully synthesized through iterative feedback between the model and experiment. The recipe generator accurately refined synthesis parameters over successive trials, demonstrating its ability to reproduce phase-pure products while minimizing trial-and-error. This approach establishes a data-driven, feedback-optimized route to accelerate synthesis design, offering a generalizable paradigm for integrating language models into experimental materials research.

Read PDF

Similar papers

Preprint Aug 2026

Synthesizing like a chemist: an iterative, feedback-driven loop for materials discovery

Most computationally predicted materials are never synthesized because conventional synthesis optimization is slow, expertise-dependent, and iterative. Here we present a closed-loop framework that automates this expert workflow by placing human tacit knowledge in the loop through a large language model (LLM) that distills synthesis knowledge from the literature, high-throughput hyperspectral imaging for rapid film evaluation, and multi-objective Bayesian optimization guided by experimental feedback. In a paired optimization campaign, LLM-assisted initialization produced more Pareto-optimal samples and higher hypervolume than a Latin hypercube sampling baseline at matched trial counts, and this advantage persisted throughout iterative optimization. We demonstrate the framework by synthesizing the previously unreported perovskite-inspired compound Rb3BiI6 as thin films and validating the optimized films by optical bandgap analysis and X-ray diffraction. The framework transforms synthesis prediction from single-shot recommendation to iterative learning, providing a generalizable strategy to accelerate automated and fully autonomous experimental materials discovery.

Fang Sheng, Steven B. Torrisi, Amanda A. Volk et al. · 0 citations
Preprint Jul 2026

Human and LLM Collaboration for Accelerated Materials Synthesis and Discovery

Although Large Language Models (LLM) and Artificial Intelligence (AI) tools have enabled a rapid increase in the generation rate of predicted materials, the rate of new materials discovery has lagged behind. This is due to the challenges associated with designing a sequence of chemical reactions to predictably produce new materials, especially in new structure types. Here, we report a study of human and LLM generated recipes for the synthesis of known and new materials. The success of the recipes is determined through in-lab experimentation, and the results are passed back to the humans and LLMs in a closed-loop process to study the effects of their collaboration. The Ruddlesden-Popper homologous series was selected for all material candidates to provide a materials phase space that is simultaneously well studied and likely to host undiscovered materials. We find that humans (H) and LLM (L) have similar success rates: 83(8)% (H) and 75(9)% (L) [known materials, round one], 17(9)% (H) and 22(10)% (L) [unknown materials, round one], 79(8)% (H) and 71(9)% (L) [known materials, round two], and 22(7)% (H) and 14(6)% (L) [unknown materials, round two]. Through this collaborative human-LLM effort, we discovered Ba3PtO5, a material with a new structural prototype that constitutes the missing 1D member of the herein reported dimensionally tunable Rock-Salt Perovskite (RSP) homologous series of the form (AX)m(ABX3)p, of which the Ruddlesden-Popper series is a subset.

G. Bassen, Wyatt Bunstine, Sarah Okandey et al. · 0 citations
Preprint Aug 2026

Strategy-first synthesis planning for complex natural products

The total synthesis of a complex molecule is among the most demanding intellectual and experimental feats in chemistry: a chemist must plan many steps ahead for how to assemble simple building blocks into an intricate target, devise backup strategies, and anticipate procedural challenges. It is also a profoundly creative activity. For half a century, efforts to automate the retrosynthetic design of natural products and other complex molecules have drawn on catalogued reactions, and the resulting tools now report near-complete success on benchmarks built from that same source. But these tools were shaped to fit benchmarked chemistry, and they falter on many natural products, the frontier of the field, whose densely functionalized, polycyclic architectures demand precisely the inventive chemistry the record contains least. Whether a machine could reasonably design such syntheses like an expert chemist does has remained unclear. Here, we show that SynthEx, an agentic framework built on large language models, plans routes to complex natural products that lie beyond the reach of conventional design algorithms. SynthEx proposes competing strategies, assembles a sequence of routine and key steps into a cohesive route, and critiques and improves its own design; the chemistry it favours is more convergent than existing tools produce, and spans a region of reaction space that catalogue-based tools cannot match. Most notably, in blinded assessments, expert chemists judged its key steps comparable to those of published human syntheses and engaged with them as genuine synthesis plans, a response algorithmic route prediction has not previously accomplished. We release routes to more than a thousand natural products as SynthAtlas, an open, interactive database, and anticipate it will become a shared resource for a collection of complex target molecules that lack existing literature routes.

Daniel P. Armstrong, X. Nguyen, Octavian Susanu et al. · 0 citations
Preprint Aug 2026

Interpretable physics-informed retrieval-augmented generation language model for end-to-end inorganic crystal synthesis planning

Synthesis planning for inorganic materials requires predicting both synthesizability and viable routes by linking microscopic thermodynamic stability with macroscopic synthesis methods, precursors, and processing conditions. Here, we develop an interpretable Physics-Informed Retrieval-Augmented Generation Language Model (PIRAG-LM) for end-to-end inorganic crystal synthesis planning. We construct a material-centered Structured Synthesis Knowledge Base (SSKB) containing route-level records for 13,820 experimentally synthesized inorganic crystals. PIRAG-LM retrieves historical precedents using chemical, structural, and thermodynamic similarity, then employs a structured LLM reasoning module to propose routes, precursors, and processing conditions and assess thermodynamic feasibility, kinetics, and accessibility. It achieves 91.4% accuracy in synthesis-method prediction, compared with 72.1% for the LLM alone, and generalizes to materials reported after the knowledge cutoff. Because the framework relies on retrieval rather than parametric memorization, its performance can be improved by expanding the SSKB without retraining the language model. Guided by PIRAG-LM, we experimentally synthesize five new compounds: BaMo0.3In0.7O2.95, BaNb0.4In0.6O2.9, Hg[B(CN)4]2, CoCo(CN)6, and SrNb2Fe2(PO4)6, via solid-state and solution routes. These results demonstrate an interpretable machine-learning approach that helps bridge computational materials discovery and experimental realization.

Wei-Jian Jiang, Ye-Nan Sha, Hui Guo et al. · 0 citations
Preprint Jul 2026

Symbolic Predicate-Guided Language Agents for Inverse Design of Perovskite Oxides

This work introduces a domain specific language (DSL)-guided strategy to improve the reasoning and design capability of LLM agents by translating natural language design rules into symbolic predicates encoded in a predefined chemistry DSL, and developed a multi-agent materials design framework.

Dong Hyeon Mok, Seoin Back, Victor Fung et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Training Chemical Plausibility-Aware Large Language Models for Single-Step Retrosynthesis

This work introduces Top-K prompting as a robust training and inference paradigm to better capture diverse, plausible reaction predictions and establishes Top-K, plausibility-aware training as a practical new direction for robust future LLM-based synthesis planning.

B. Zagribelnyy, Ivan D. Ilin, N. Bondarev et al. · 0 citations