Skip to content
Open access

A Text-Based Refactoring Dataset for Multi-Type Prediction Using Code Snippets and Software Metrics

2026 · IEEE Access · Vol 14, pp. 140093-140113 · 0 citations · 31 references

Abstract

Refactoring is a disciplined process of improving the internal structure of software without changing its external behavior. Empirical studies have shown that refactoring contributes to maintainability and code quality; however, existing machine learning approaches for refactoring prediction rely predominantly on structural and process metrics, without incorporating the actual source code content. This limits the applicability of deep learning and representation learning methods that require direct access to code. This study introduces a text-based refactoring dataset constructed from 51 open-source Java repositories, comprising 159,628 instances spanning 18 refactoring types across class, method, and variable levels. Each instance combines structural and process metrics inherited from the RefactoringML corpus with before-refactoring source code snippets extracted via an ANTLR4-based AST parsing pipeline. To characterize the difficulty of refactoring prediction tasks and demonstrate the utility of the dataset, we evaluate three classification tasks using CodeBERT-based code embeddings and traditional software metrics. Results show that predicting the abstraction level (Class, Method, or Variable) achieves a macro F1 of 0.7689 using a hybrid feature set, while fine-grained prediction of 18 specific refactoring operations remains challenging (best macro F $1=0.2512$ ). The dataset is designed to support future research on ML-based and LLM-based refactoring prediction by providing a resource that combines structural metrics with textual code representations.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.