Skip to content
Book Open access

Semantic-Enhanced Automatic Refinement of Architecture Recovery Results Using LLMs

Apr 2026 · Proceedings of the 2026 IEEE/ACM 48th International Conference on Software Engineering · pp. 1687-1699 · 1 citation · 84 references
Computer Science

TL;DR

SemRef, a framework that combines LLMs with dependency analysis to automatically refine architectures recovered by existing architecture recovery tools, is introduced, which improves accuracy across various metrics and enhances both the accuracy and the comprehension of recovered architectures.

Abstract

Understanding the architecture is crucial for effectively maintaining and managing large software systems. However, discrepancies often exist between the designed and implemented architectures, which can pose significant risks. To identify these discrepancies, architects need to extract the architecture from the system implementation, which is both time-consuming and error-prone. To simplify this procedure, many automatic architecture recovery techniques have been developed. Yet, their accuracy is often limited. Architects must still invest significant effort in refining recovery results to ensure they accurately reflect the implemented architecture. To reduce such manual effort, we introduce SemRef, a framework that combines LLMs with dependency analysis to automatically refine architectures recovered by existing architecture recovery tools. By leveraging the LLM’s semantic understanding capabilities and integrating structural dependencies, SemRef enhances both the accuracy and the comprehension of recovered architectures. To evaluate SemRef, we tested on 9 projects with published ground-truth architectures and 10 state-of-the-art architecture recovery tools. 5 commonly used metrics are adopted to evaluate the effectiveness of SemRef. The results show that SemRef improves accuracy across various metrics, with normalized gains ranges from 17.72% to 43.35%. Specifically, for MoJoFM and a2aadj metrics, SemRef achieves relative improvements of 118.57% and 100.41%, respectively. Moreover, SemRef is highly scalable. It maintains stable performance across projects ranging from thousands to trillions of lines of code with the cost scale linearly with project size. Further, we test SemRef on various LLMs to demonstrate its generalizability across different models. Beyond improving accuracy, the integration of LLMs enables SemRef to provide a structured module hierarchy and hierarchical module summaries, which further enhance the comprehensibility of recovered architectures.

Read PDF

Similar papers

Open access Aug 2026

Architecting Reliable Knowledge Retrieval Systems Using Large Language Models

A literature-based architectural framework for reliable knowledge retrieval systems that separates external knowledge management from LLM-based reasoning and generation is developed and indicates that reliable LLM deployment should be treated as an end-to-end architectural problem rather than solely a model-performance...

Bharat Kumar Reddy Karumuri · 0 citations
Conference Aug 2026

Toward Cost-Efficient Automated Requirements Traceability with Large Language Models

Automated requirements traceability is a critical activity in software engineering, supporting impact analysis, verification and validation, and regulatory compliance. As modern software-intensive systems grow in scale and complexity, maintaining trace links manually becomes increasingly impractical. Recent work has sh...

Nouf Alturayeif, Jameleddine Hassine, Irfan Ahmad · 0 citations
#software testing Preprint Sep 2026

Relationally Guided Use Case Modeling with LLMs

The proposed FlowGen uses LLM-based Semantic Information Processing (SIP) to extract semantic elements, constructs a Semantic Relational Graph (SRG) encoded by an enhanced R-GAT for basic flow generation (BFGen), and further supports branch point prediction through BPP and branch-conditioned alternative flow generation...

Guang-Yu Wang, Bangqi Li, Ji Wu et al. · 0 citations
Preprint Aug 2026

How Well Do LLMs Generate Taxonomies in the SE Domain? A Multi-perspective Evaluation Framework

Taxonomies provide a shared conceptual framework for organizing heterogeneous observations in software engineering (SE) research. Manually constructing such taxonomies is labor-intensive and requires annotators with expertise in the SE domain. While advances in Large Language Models (LLMs) have led to the emergence of...

Sota Nakashima, Yuta Ishimoto, Masanari Kondo et al. · 0 citations
Book Open access Oct 2026

A Transformation-Based Benchmark for Evaluating the Robustness of LLMs in Generating OCL

Large Language Models (LLMs) have shown promising performance in generating Object Constraint Language (OCL) constraints from natural language specifications. However, existing evaluations rely on publicly available UML models, which may overestimate generalization due to potential data leakage and reliance on recurrin...

Hamza Attarwala, Moataz Chouchen, Omar Alam et al. · 0 citations
Preprint Aug 2026

An Empirical Study on the Impact of Normalized Use-Case Specifications on Traceability

Traceability link recovery between requirements and source code is vital for software quality assurance and evolution analysis. Although automated traceability techniques have advanced greatly, the large semantic gap between vague natural-language requirements and precise source code still hinders accurate link recover...

Luoyuan Shi, Yuan-Zhao Zhai, Da-Wei Feng et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.