Beyond Single-run Correctness: Nondeterminism-aware Evaluation of LLM-based Model Transformations
Abstract
Model transformation is a core model-driven engineering (MDE) operation in which reproducibility is expected: under fixed metamodels, source model, and transformation rules, a deterministic engine should produce a stable target model. Large Language Models (LLMs) are increasingly explored for MDE tasks, but evaluations often focus on whether an acceptable artifact can be produced once. For deterministic and quasi-deterministic MDE workflows, this single-run correctness view is necessary but insufficient. This paper introduces nondeterminism-aware evaluation for LLM-based MDE. Using model-to-model transformation as a stress case, we compare LLM-generated target models against deterministic ATL references and across repeated executions. Our exploratory evaluation covers five ATL Zoo transformation scenarios, four prompt configurations, three LLMs, and ten executions per transformation -scenario–configuration–LLM combination. We analyze reference deviation, inter-run variation, and morphological differences. The results show that transformation explicitness does not guarantee convergence. Even when the complete ATL transformation is provided, 13 out of 15 transformation-scenario–LLM combinations show non-zero mean and median deviation from the ATL reference, and 10 out of 15 exhibit non-zero inter-run variation. We further observe stable but not ATL-equivalent behavior and cases where a single exact match coexists with non-zero median deviation. These findings suggest that LLM-based MDE evaluation should treat correctness, reproducibility, and variation meaning as distinct dimensions.