Aug 2026· Proceedings of the ACM/IEEE 29th International Conference on Model Driven Engineering Languages and Systems· pp. 200-211· 1 citation· 65 references
Computer Science
TL;DR
This paper proposes an automated approach to extract domain models from source code using lightweight, locally deployable LLMs and achieves high F1-scores on a dataset of ten projects, each comprising a curated domain model and its corresponding implementation, while remaining fully executable on locally deployable LLMs.
Abstract
Large language models (LLMs) have recently shown strong capabilities for code understanding, making them promising for reverse engineering domain models from source code. However, state-of-the-art proprietary LLMs cannot be used in many industrial contexts due to privacy and confidentiality constraints, while compact open-source LLMs that can run locally are limited by their context window and cannot process large code bases directly. In this paper, we propose an automated approach to extract domain models from source code using lightweight, locally deployable LLMs. Our method combines structural and semantic heuristics with iterative LLM-based reasoning to overcome context limitations. By progressively analyzing ranked subsets of code elements, the approach identifies domain concepts and refines domain boundaries without requiring full-system context. Our approach achieves high F1-scores on a dataset of ten projects, each comprising a curated domain model and its corresponding implementation, while remaining fully executable on locally deployable LLMs. This makes it particularly suitable for reverse engineering tasks in privacy-sensitive industrial environments.
This paper proposes an automated approach to extract domain models from source code using lightweight, locally deployable LLMs and achieves high F1-scores on a dataset of ten projects, each comprising a curated domain model and its corresponding implementation, while remaining fully executable on locally deployable LLM...
Alessandra Mancas, Mounir Ammam, Hyacinth Ali et al.· 0 citations
Large Language Models (LLMs) are increasingly used for software vulnerability detection, but their performance depends on how source code is represented in the input. Most prompting approaches use source code in its original form, while some works propose the use of structured representations. Abstract Syntax Trees (AS...
In an era where software development is deeply tied with Large Language Models, does Model Driven Engineering (MDE) still make sense? This raises the question of the extent to which MDE can be successfully combined with an LLM approach to address the downsides of each approach separately. In this paper, we try to answe...
Marco Calamo, Massimo Mecella, Monique Snoeck· 0 citations
Automated requirements traceability is a critical activity in software engineering, supporting impact analysis, verification and validation, and regulatory compliance. As modern software-intensive systems grow in scale and complexity, maintaining trace links manually becomes increasingly impractical. Recent work has sh...
Large language models (LLMs) are increasingly used for model-to-code generation, but they weaken one of the central properties of model-driven engineering: traceability between design models and generated implementation artefacts. Existing approaches either rely on deterministic transformations, where trace links can b...
Marc North, Nelly Bencomo, Amir Atapour-Abarghouei· Proceedings of the ACM/IEEE...· 0 citations
Large language models (LLMs) are increasingly used in software pipelines, raising concerns about harmful behaviors in security-critical domains. Existing safety evaluations predominantly probe models with single prompts or short interactions, and therefore do not capture how safety behaves under multi-step workflows wh...
Lu Yan, Zhuo Zhang, Xiang-Zhe Xu et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.