From Scientific Copilots to Tool-Grounded Autonomy: AI Agents in Simulation-Driven Materials Discovery
Artificial intelligence (AI) agents and large language model (LLM) agents are beginning to move materials discovery beyond isolated prediction tasks and toward tool-grounded workflows that can retrieve prior knowledge, configure simulations, launch calculations, inspect outputs, and decide what to do next. However, adjacent reviews on materials informatics, self-driving laboratories, natural-language processing in materials science, and autonomous chemistry have not isolated simulation-driven materials workflows as a distinct evidence base. This review addresses that gap through PRISMA-guided searches in Scopus (8 May 2026) and Web of Science (15 June 2026) for English-language journal articles published between 2022 and 2026. The combined search returned 232 records; 27 full texts were assessed and 26 studies were included in the final qualitative synthesis after one full-text exclusion. No eligible study was published in 2022 or 2023, indicating that the field emerged only in 2024 and expanded rapidly in 2025–2026. Catalysis and adsorption tasks (n = 6) and alloy design or evaluation (n = 5) dominated the corpus, while specialized multi-agent architectures were the most common pattern (n = 13). Across the included studies, agentic reasoning was most often coupled to workflow orchestration or integration tools, molecular-dynamics or atomistic simulation environments, and materials-data or machine learning screening pipelines; public repositories or archival artifacts were reported in 18 of 26 studies, experimental validation in six, and robotic closed-loop execution in only one study. The strongest evidence came from workflows that grounded language model decisions in simulators, structured databases, or experimentally verifiable outputs rather than in free-form text alone. This review therefore establishes AI agent workflow orchestration as a distinct analytical category within materials discovery and identifies the reporting, validation, and reproducibility conditions required for these systems to function as credible scientific infrastructure rather than as conversational demonstrations.