This survey provides the first systematic synthesis of LLM-based diagram modelling research, highlighting needs for standardised benchmarks, stronger evaluation protocols, broader diagram coverage, and techniques for improving semantic reliability and multi-view consistency.
Abstract
Large language models (LLMs) are increasingly applied to diagram-based software and data modelling. Among various modelling notations, UML and entity-relationship (ER) diagrams are the most widely adopted for software modelling and data modelling, respectively. Recent literature has investigated various applications of LLMs in diagram modelling; however, their effectiveness and limitations have not been extensively discussed. This systematic literature review analyses 64 studies published between 2023 and 2025, examining diagram coverage, modelling tasks, technical approaches, evaluation practices, and limitations. Our findings reveal significant concentration patterns and gaps. UML-based software modelling strongly dominates, with class diagrams receiving the most attention whilst behavioural diagrams and data modelling remain underrepresented. Diagram construction from natural language is the primary focus, with limited work on transformation, quality assurance, and consistency checking. GPT-based models are heavily prevalent, raising concerns about reproducibility and vendor dependence. Evaluation practices are heterogeneous, employing diverse metrics and custom datasets with limited benchmark reuse and inconsistent reporting of robustness and statistical significance. Common limitations include semantic inaccuracies, hallucinated diagram elements, sensitivity to prompt formulation, and reproducibility constraints. This survey provides the first systematic synthesis of LLM-based diagram modelling research, highlighting needs for standardised benchmarks, stronger evaluation protocols, broader diagram coverage, and techniques for improving semantic reliability and multi-view consistency.
A systematic literature review of empirical studies on UML SDs in computing education, with a focus on notation, consistency, completeness, and model quality shows that research on comprehension represents the primary focus of literature, followed by studies addressing model quality, while relatively few studies focus on methods and tools.
Sohail Alhazmi· Annual Conference on Innovat...· 0 citations
This paper investigates the reliability of LLMs in evaluating UML diagrams generated through reverse engineering processes (source code) and asks: do LLM assessments align with those of human experts?
Olena Chebanyuk, Carles Sierra· Proceedings of the 21st Inte...· 0 citations
The automated generation of business process models from natural language descriptions has recently attracted growing attention at the intersection of Business Process Management (BPM), Natural Language Processing (NLP), and Large Language Models (LLMs). This paper presents a Systematic Literature Review (SLR) on the current state of research in this emerging field. Following the guidelines of Kitchenham et al. and the PRISMA framework, 29 studies published between January 1st, 2023 and March 10th, 2026 were identified, selected, and analyzed. The review addresses two research questions focusing on the applied methodological approaches, the used LLMs, the employed modeling languages, as well as the evaluation strategies, challenges, and limitations reported in the literature. The results show a clear shift from traditional NLP-based techniques toward LLM-only and hybrid approaches. OpenAI’s GPT family, especially GPT-4 and its variants, dominates the field, while BPMN is by far the most frequently used target process modeling language. Furthermore, existing studies evaluate automated process model generation primarily through output-focused methods, such as quantitative metrics, expert reviews, and comparisons with alternative or human-created process models. At the same time, the reviewed studies reveal important challenges, including the continued need for human involvement and the output quality. Overall, current approaches show strong potential, but they still act more as intelligent assistants than as fully autonomous process modelers.
L. F. Hörner, Maximilian Möller, Manfred Reichert· IEEE Access· 0 citations
This work presents the first cross-task empirical evaluation of LLMs spanning five RE-related activities, as well as replication materials supporting reproducibility, and a broader understanding of the capabilities, limitations, and practical readiness of current LLMs for RE.
Jacek Dabrowski, Manjeshwar Aniruddh Mallya, Alessio Ferrari et al.· 0 citations
The results show that the applied LLM can reliably detect structural and semantic differences between formal business process models using Business Process Model and Notation, while distinguishing them from acceptable variations, demonstrating strong potential for automated model validation.
Christian Bennoit, S. Zamani, Tobias Greff· Process Science· 0 citations