This review surveys how Large Language Models are adding semantic interfaces, code generation, and tool orchestration to established numerical nanophotonic workflows, and looks ahead to the next generation of multimodal foundation models with physical perception capabilities.
Abstract
Metasurfaces have revolutionized the development of photonic devices by enabling unprecedented precision in light manipulation. However, their design processes are often constrained by computationally expensive simulations and complex high-dimensional design spaces. Although deep learning has accelerated the design process by serving as a surrogate model, it remains constrained by task-specific architectures and lacks universal reasoning capabilities. This review surveys how Large Language Models (LLMs) are adding semantic interfaces, code generation, and tool orchestration to established numerical nanophotonic workflows. We first outline the development from classical neural networks to transformer-based models and their applications in nanophotonic design. We then review the emergence of LLM-related methods in nanophotonics and organize them into two operational modes: surrogate models that treat structure-spectrum mapping as a language task, and agentic systems that have been demonstrated to generate code, orchestrate selected simulation steps, and support closed-loop optimization. Furthermore, to identify future cross-disciplinary opportunities, we briefly explore applications of LLMs in research fields such as materials science and wireless communications. This review concludes by looking ahead to the next generation of multimodal foundation models with physical perception capabilities. In this vision, artificial intelligence is evolving from passive tools into active collaborators, participating in autonomous scientific discovery.
This review traces the development of the field from classical machine learning and deep learning to generative models, transfer learning, transformers, and emerging foundation models, and introduces major nanophotonic platforms.
Chaobin Yang, Xueqing Liu, Yiqun Fu et al.· 0 citations
Metasurface design increasingly requires fast models that can operate across structurally distinct device families, rather than retraining a separate surrogate for every geometry class. Conventional neural network surrogates often depend on fixed-dimensional descriptors, family-specific output formats, and repeated architecture tuning, which limits their scalability across heterogeneous meta-atoms. Here, we present a unified large language model (LLM) workflow for multi-family metasurface modeling and inverse-design. Geometries, design parameters, and optical response channels were converted into a shared instruction-following text format and used to fine-tune Gemma-2-9B across 8 metasurface families. Compared with single-family baselines, the joint model simultaneously predicted the optical responses of all metasurface families while reducing the MSE for each family by an average of 56.5%. The same representation was also used for inverse design. These results show that a shared sequence-based LLM interface can provide a practical route to cross-family metasurface design while reducing the need for task-specific surrogate architectures.
Huanshu Zhang, Lei Kang, Yu-Yan Chen et al.· 0 citations
ABSTRACT Metamaterials for light manipulation at subwavelength scales face significant design challenges due to their complex and sophisticated structures, leading to the emergence of deep learning as a powerful tool to streamline their design process. However, existing deep learning‐based inverse design methods fall short in the design of reconfigurable metamaterials (RMMs), whose optical characteristics switch between different states upon external stimuli. To address this challenge, CoSP, an intelligent inverse design method for RMMs based on a contrastive pretrained large language model (LLM), is proposed. By performing contrastive pretraining on multi‐state spectra, a well‐trained spectrum encoder is obtained and coupled to a GPT‐style decoder trained end‐to‐end from scratch. Equipped with the preservation of linguistic capabilities, CoSP is capable of describing material structures with target optical properties in natural language. Numerical experiments demonstrate that CoSP can design RMM structures for multi‐state, multi‐band optical responses, showing great potential in versatile applications such as thermal management, optical computation, and telecommunications.
Shujie Yang, Yuqi Zhang, Xuzhe Zhao et al.· Advancement of science· 0 citations
Results demonstrate that an LLM agent can assume key aspects of expert decision-making in photonic inverse design while maintaining physical fidelity and fabrication feasibility, providing a scalable route towards autonomous design of complex integrated photonic systems.
Faqian Chong, Yulun Wu, Shilong Li et al.· 0 citations
Significance Computational modeling and the use of simulation software tools are essential for biomedical optics research. Designing effective simulations often requires in-depth understanding of the underlying physical problems and proper configuration of the software settings, which often constitute key barriers for novice users including students. The rapid emergence of large language models (LLMs) offers new opportunities for natural-language-based interaction, but integrating them with technical software remains challenging because of their limited output reproducibility. Overcoming these limitations would allow more intuitive, efficient, and reproducible interaction between scientists and scientific software. Aim We investigate the use of LLMs in quantitative biophotonics simulation tools, with a goal of enabling novice users to build complex photon simulations using intuitive natural-language-based problem descriptions. Approach We have explored prompt engineering strategies that enable LLMs to bridge the gap between natural language descriptions and advanced simulation software by constraining LLM outputs using a data schema (i.e., format) and a modular component architecture, followed by deterministic validation to ensure correctness and reproducibility of the outputs. Results Using Monte Carlo eXtreme (MCX) – a widely used photon transport simulator – as an example, we showcase the capability of the proposed framework to convert user descriptions to structured simulation inputs. Benchmarked using 33 diverse natural language simulation descriptions, our LLM interface, MCX-LLM, achieves 98% accuracy and 99% repeatability, with an average processing time of 8.96 seconds per prompt. The framework also successfully handles various linguistic styles and diverse simulation settings, achieving a 100% success rate on 20 unconstrained real-world prompts. With only minor adjustments, our LLM interface also produces valid inputs for a finite-element-based diffusion solver to demonstrate generality towards other optical simulators. Conclusions By combining LLMs’ capability for textual data comprehension with structured constraints, this work provides a pathway to making complex scientific tools accessible while ensuring the reliability and technical correctness required for rigorous scientific research. MCX-LLM has been integrated with MCX Cloud accessible at https://mcx.space/cloud.
PICopilot is introduced, the first large language model (LLM)-based agentic framework that assists in PIC design via automated design script generation from natural language instructions, achieving a high success rate and reliability.
Xiaohan Jiang, Zeyu Li, Wei Zhang et al.· 0 citations
A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.
MIT News · Artificial Intelligence· news.mit.eduAug 24, 2026