A Comprehensive Evaluation of LLMs for Design Pattern Recognition in Software Systems
Abstract
Design patterns are fundamental to high-quality software, yet their documentation is often incomplete or absent, making automatic design pattern recognition (DPR) essential for reverse engineering, refactoring, and long-term program comprehension. Traditional DPR methods rely on hand-crafted rules or extensive feature engineering, limiting scalability and generalisability. This study presents a lightweight, rule-free framework that leverages pre-trained large language model (LLM) embeddings combined with a simple $k$-Nearest Neighbours $(k$-NN) classifier. Five state-of-the-art embeddings CodeBERT + $k$-NN, CodeT5 + $k$-NN, TinyLlama-1.1B-Chat-v1.0, RoBERTa + $k$-NN, and GraphCodeBERT + $k$-NN were evaluated on 93 Java programs from the P-MARt repository across six classes (AbstractFactory, Builder, FactoryMethod, non-implemented DP, Prototype, and Singleton). GraphCodeBERT + $k$-NN and CodeT5 + $k$-NN achieved the highest mean $F_{\mathrm{1}}$-scores $(\text{0. 9 6 9}$ and $\text{0. 9 5 3})$, closely followed by RoBERTa $\boldsymbol{+} \boldsymbol{k}$-NN (0.941), while TinyLlama performed substantially worse (0.580). Analysis of classification consistency further reveals that semantics, syntax, implementation style, and contextual clarity are the primary drivers of reliable detection. These findings establish a practical, reproducible baseline for LLM-enhanced DPR and highlight its potential to transform software maintenance and architectural recovery practices.