Skip to content
Open access

GenAI-augmented ensemble learning framework for phishing URL detection using lexical, host-based and sequential features

Aug 2026 · Международный Журнал Теоретических и Прикладных Вопросов Цифровых Технологий · 0 citations

TL;DR

Experimental results show that the proposed stacking ensemble consistently achieves the best overall performance, while a moderate augmentation ratio of 25% provides the highest robustness under temporal and cross-dataset evaluation; higher augmentation levels reduce performance.

Abstract

Phishing remains one of the most persistent cybersecurity threats, and URLs are often the earliest observable indicator of an attack. Although existing phishing URL detection methods achieve high accuracy using lexical, host-based, deep learning, and ensemble approaches, many rely on static evaluation settings that overlook temporal drift, domain leakage, cross-dataset bias, and adversarial URL evolution. This study proposes a leakage-aware phishing URL detection framework that combines lexical, host-based, and sequential URL representations with GenAI-augmented ensemble learning. A safety-filtered synthetic URL generation module produces realistic phishing patterns, including brand impersonation, typosquatting, excessive subdomains, homoglyph variants, suspicious paths, and malicious query structures, without generating live malicious domains. Classical machine learning models, deep sequence networks, transformer-lite models, and ensemble methods are evaluated using random, temporal, domain-disjoint, and cross-dataset splits. The impact of synthetic augmentation ratios (0%, 10%, 25%, and 50%) is assessed using F1-score, ROC-AUC, PR-AUC, false positive rate, and false negative rate. Experimental results show that the proposed stacking ensemble consistently achieves the best overall performance, while a moderate augmentation ratio of 25% provides the highest robustness under temporal and cross-dataset evaluation; higher augmentation levels reduce performance. The proposed framework offers a reproducible and leakage-aware benchmark for evaluating whether GenAI-based data augmentation and multi-signal ensemble learning improve resilience against evolving phishing URL attacks.

Read PDF

Similar papers

Open access Aug 2026

Multi-Source Generalization-Aware Phishing URL Detection Using Calibrated Stacked Ensemble and False-Positive Control

GAFPNet (Generalization-Aware and False-Positive Controlled Framework Network), a five-module stacked-ensemble framework, is introduced as a generalization-aware and deployment-oriented phishing URL classifier for real-time filtering applications.

Mohammed Elias Basha S., M. N. · 0 citations
Open access Aug 2026

A feature-enriched deep learning based ensemble framework for robust phishing URL detection

A feature-driven framework for phishing Uniform Resource Locator (URL) detection is introduced, emphasizing the design and evaluation of enhanced feature representations and highlighting that performance gains are primarily driven by feature design rather than model complexity.

Deniz Kaya, Murat Osmanoğlu · 0 citations
Conference Jul 2026

Multi Model Approach for Phishing Website Detection using ML and DL Techniques

In the technology era, Phishing has continued to be a great challenge within the cybersecurity and web security landscape. This involves exploiting human trust on any online services and subtle technical flaws. This is to gather credentials, financial data, and sensitive information across diverse online platforms and various users. Traditional defenses like static blacklists, signature-based filters and simple detection rules are limited by slow update cycles and an inability to capture subtle syntactic and behavioral cues. To address these shortcomings, we propose a hybrid detection framework that fuses classical supervised machine-learning classifiers (e.g., Logistic Regression, SVM, Random Forest, XGBoost) with sequence-aware deep learning (LSTM) to jointly model lexical, structural, syntactic, and behavioral features extracted from URLs and webpage metadata. This combined approach leverages the interpretability and stability of ML models alongside the pattern-learning strength of LSTMs to detect both known and zero-day phishing attempts, produce calibrated confidence scores and deliver comprehensive reports via a real-time web interface resulting in a robust, transparent, and operationally useful solution for strengthening web security.

M. Yaswanth, Pathan Basheer Khan, Dhulipalla Naga Harish et al. · 0 citations
Open access Jul 2026

ELEVATING PHISHING DETECTION PERFORMANCE WITH MACHINE LEARNING AND DEEP LEARNING-ENABLED FEATURE SELECTION

An intelligent phishing detection framework that integrates machine learning, deep learning-enabled feature selection, multi-source phishing feature extraction, hybrid classification, and real-time risk assessment is proposed that can achieve higher detection accuracy, precision, recall, F1-score, and lower response latency than blacklist-based, conventional machine learning, and standalone deep learning approaches.

P. Paul, Bharath Bhushan, Bandameedi Sai Charan et al. · 0 citations
Open access Jul 2026

Large Language Models for phishing URL detection: A comparative study of LLaMA-3, GEMMA-7B, and traditional Machine Learning approaches

The proposed LLM-based framework offers a promising approach for improving phishing detection and strengthening modern cybersecurity defenses and suggests that transformer-based models can effectively identify deceptive domain structures, abnormal URL patterns, and obfuscation techniques.

L. Eliyan, M. Alshraideh, Bayan Alfayoumi · 0 citations