Skip to content
Review Open access

Curriculum Learning for Fine-Tuning Code-Review Language Models: A Case Study in Diagnosing and Correcting Fabrication in Vulnerability Explanation

Sep 2026 · International Journal for Research in Applied Science and Engineering Technology · Vol 14, pp. 33-40 · 0 citations

TL;DR

Results show that target design and manual inspection are important when evaluating generated vulnerability explanations, and that negative results, like the target-length test here, are worth reporting alongside positive ones.

Abstract

Automated code-review systems can identify potential defects, but they often provide limited explanations for developers. This study investigates a curriculum-based approach for fine-tuning CodeT5+ 770M to generate code-review and vulnerability explanations for Python and JavaScript. Training was divided into three stages: code understanding, semantic review, and vulnerability explanation. Rehearsal examples from earlier stages were included during later training to reduce forgetting. We also compared this staged approach with an earlier mixed-phase setup that trained on all tasks together, to see whether staging changed how the training process behaved. During Stage 3, we found that using case-specific CVE descriptions as targets led the model to generate unsupported package names, versions, and repository references. We replaced these targets with CWE-level descriptions and retrained the model. In a manual review of 20 outputs, the revised model did not produce the same type of unsupported specific details. BLEU increased from 7.54 to 34.43, and ROUGE-1 increased from 0.2765 to 0.4180. However, the revised model still confused CWE categories: exact category agreement was 27.3% after excluding samples with incomplete ground-truth labels, and Cross-Site Scripting was predicted more often than its true frequency. We also tested whether increasing the Stage 2 target length from 192 to 512 tokens improved later Stage 3 performance. Across six epochs, the average validation-loss difference was +0.0008, and the final Stage 3 metrics showed no meaningful change. These results show that target design and manual inspection are important when evaluating generated vulnerability explanations, and that negative results, like the target-length test here, are worth reporting alongside positive ones.

Read PDF

Similar papers

Open access

Evaluation and Distillation of Source Code Generation Tasks by Large Language Models

Two novel contributions are introduced: CodeEval and CodeQual, an open-source execution framework that provides researchers with a ready-to-use evaluation pipeline for evaluating and improving LLMs in software engineering contexts, encompassing both functional correctness assessment and subjective code quality evaluati...

Danny Brahman · 0 citations
Open access Aug 2026

Do Pre-Trained Code Models Add Value Beyond Software Metrics in Class-Level Defect Prediction? An Empirical Study of Input Coverage, Long-Code Aggregation, and Cross-Version Generalization

Findings indicate that pre-trained code models should be evaluated with explicit input-coverage reporting, chronological validation, strong metric baselines, and incremental-value testing.

Musaad Alzahrani · 0 citations
#software testing Book Open access Oct 2026

Curated Semantic Mutants: Multi-purpose Artifacts for Grading and Hinting Student Test Suites

A human-LLM workflow that pairs each curated semantic mutant with an instructor-approved seed phrase for an on-demand LLM expansion, which shows that the curated semantic-mutant set contains a fraction of the mutants a traditional mutation engine produces.

Rebecca Williams Earle, Jonathan Bell · 0 citations
2026

Generative vs. Discriminative? An Empirical Study on Code Understanding Classification Tasks

Automated code understanding is crucial for software reliability and maintainability. Encoder-only pretrained models excel in code classification tasks such as vulnerability detection, cross-language clone detection, and exception classification due to their bidirectional context awareness. However, the dominant “pre-t...

Yu-Guo Liu, Yu-Hua Ma, Rong-Cun Wang et al. · 0 citations
Preprint Sep 2026

SpecCoder: Specification-Aware Code Generation with Curriculum Dual-Task Reinforcement Learning

Large language models (LLMs) have made substantial progress in code generation but still struggle with challenging programming tasks that require understanding rich natural language requirements. These requirements often specify problem goals, input/output formats, constraints, examples, and edge cases. Overlooking eve...

Yi-Xuan Li, Min Huang, Jia-Jing Wang et al. · 0 citations
Review Oct 2026

Correctness, Convergence, and AI-Generated Code Detection: A Longitudinal Study of Student and Large Language Model Code in Introductory Programming

Large language models can generate plausible solutions to programming assignments, making it tempting to detect their use by matching student code against a reference bank of generated solutions. Yet similar code can also arise when an assignment admits only a few natural implementations, which leaves open what a match...

Runlong Ye, Jing Fan, Angela Zavaleta Bernuy et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.