Skip to content
Open access

H-LLM-IM: An Adaptive Imbalance-Aware Hybrid Framework Using LLM-Based Code Embeddings for Software Vulnerability Detection

Aug 2026 · Süleyman Demirel Üniversitesi Fen Bilimleri Enstitüsü Dergisi · Vol 30, pp. 350-369 · 0 citations · 23 references

TL;DR

The results demonstrate that the proposed adaptive framework substantially improves minority-class vulnerability detection, achieving up to a 2.5-fold increase in F1-score and up to 83% improvement in MCC compared to the no-imbalance baseline.

Abstract

Software vulnerabilities remain a critical threat to modern software systems, while existing detection approaches often suffer from high computational cost, limited scalability, and severe class imbalance in real-world datasets. To address these challenges, this study proposes H-LLM-IM, an adaptive imbalance-aware hybrid framework for efficient software vulnerability detection. The proposed framework leverages semantic code embeddings extracted from a pre-trained code language model (CodeBERT) and integrates them with lightweight machine learning classifiers, thereby avoiding expensive fine-tuning of large language models.A key contribution of H-LLM-IM is an adaptive imbalance-aware learning mechanism that dynamically regulates imbalance mitigation intensity through controlled oversampling and adaptive reweighting based on minority-class performance feedback. Extensive experiments conducted on the Big-Vul benchmark dataset evaluate four classifiers (Logistic Regression, SVM, Random Forest, and XGBoost) under multiple imbalance-handling scenarios, including static and adaptive strategies.The results demonstrate that the proposed adaptive framework substantially improves minority-class vulnerability detection, achieving up to a 2.5-fold increase in F1-score and up to 83% improvement in MCC compared to the no-imbalance baseline. Importantly, these performance gains are obtained while maintaining practical training time and controlled memory growth. In particular, Logistic Regression and XGBoost exhibit the most favorable performance–efficiency trade-off, highlighting the scalability and practical applicability of H-LLM-IM for large-scale vulnerability analysis under severe class imbalance.

Read PDF

Similar papers

Open access 2026

Malicious Repository Detection Using Transformer-Based Models: A Metadata-Driven Approach

A metadata-driven detection framework that fine-tunes DistilBERT on repository metadata text, specifically the description, README, and topics, and formulates detection as a binary classification task between malware and benign repositories, suggesting that the framework is promising for threshold-based screening scena...

Rabeaa Mouty, M. Abdullah-Al-Wadud · 0 citations
Open access Aug 2026

Enhancing Vulnerability Detection Precision through Ensemble Learning with Large Language Models

The results show that the ensemble techniques are a practical approach to boost the precision of LLMs in the detection of vulnerabilities and suggest that ensemble methods offer great potential in the advancement of software security analysis.

H. Al-Ofeishat, Azhar Hussain, M. Faheem et al. · 0 citations
2026

LLMKernelBench: Benchmarking Large Language Models on Software Vulnerability Detection in Linux Kernel

Large language models (LLMs) demonstrate strong capabilities in code-related tasks, however their effectiveness in software vulnerability detection (SVD) remains poorly understood due to inadequate evaluation frameworks. Existing benchmarks suffer from training data contamination, isolated function evaluation without c...

Arastoo Zibaeirad, Rodrigo Pato Nogueira, Marco Vieira · 0 citations
Open access Aug 2026

Multilingual Source Code Vulnerability Detection Using Deep Learning: A Semantic Representation and Transfer Learning Approach

This work presents a deep learning approach for multilingual vulnerability detection that emphasizes semantic transfer rather than architectural complexity and suggests that stabilizing semantic representations during transfer is key to improving generalization while maintaining practical efficiency under moderate comp...

Tuan Nguyen Kim, Nin Ho Le Viet, Chieu Ta Quang · 0 citations
Open access Aug 2026

Optimizing Multiclass Android Malware Family Classification Using SMOTE-Tomek Links and XGBoost

The increasing sophistication of Android malware attacks has created significant challenges for accurate malware family classification, particularly under highly imbalanced data distributions where minority malware families are frequently misclassified. This study presents a robust multiclass Android malware family cla...

Ali Nur Ikhsan, Adam Prayogo Kuncoro, Debby Ummul Hidayah et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Comparative Evaluation of Static Embedding Models for HTTP Request Anomaly Detection

Web applications are increasingly targeted by cyberattacks that exploit HTTP requests to evade security mechanisms. Traditional web application firewalls (WAFs) rely on rule-based approaches that often exhibit high false positive rates and limited adaptability. Recent studies have explored machine learning techniques a...

A. Riverol, Gustavo Betarte, R. Martínez et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.