Aug 2026· Международный Журнал Теоретических и Прикладных Вопросов Цифровых Технологий· 0 citations
TL;DR
Experimental results show that the proposed stacking ensemble consistently achieves the best overall performance, while a moderate augmentation ratio of 25% provides the highest robustness under temporal and cross-dataset evaluation; higher augmentation levels reduce performance.
Abstract
Phishing remains one of the most persistent cybersecurity threats, and URLs are often the earliest observable indicator of an attack. Although existing phishing URL detection methods achieve high accuracy using lexical, host-based, deep learning, and ensemble approaches, many rely on static evaluation settings that overlook temporal drift, domain leakage, cross-dataset bias, and adversarial URL evolution. This study proposes a leakage-aware phishing URL detection framework that combines lexical, host-based, and sequential URL representations with GenAI-augmented ensemble learning. A safety-filtered synthetic URL generation module produces realistic phishing patterns, including brand impersonation, typosquatting, excessive subdomains, homoglyph variants, suspicious paths, and malicious query structures, without generating live malicious domains. Classical machine learning models, deep sequence networks, transformer-lite models, and ensemble methods are evaluated using random, temporal, domain-disjoint, and cross-dataset splits. The impact of synthetic augmentation ratios (0%, 10%, 25%, and 50%) is assessed using F1-score, ROC-AUC, PR-AUC, false positive rate, and false negative rate. Experimental results show that the proposed stacking ensemble consistently achieves the best overall performance, while a moderate augmentation ratio of 25% provides the highest robustness under temporal and cross-dataset evaluation; higher augmentation levels reduce performance. The proposed framework offers a reproducible and leakage-aware benchmark for evaluating whether GenAI-based data augmentation and multi-signal ensemble learning improve resilience against evolving phishing URL attacks.
GAFPNet (Generalization-Aware and False-Positive Controlled Framework Network), a five-module stacked-ensemble framework, is introduced as a generalization-aware and deployment-oriented phishing URL classifier for real-time filtering applications.
Mohammed Elias Basha S., M. N.· Journal of Trends in Compute...· 0 citations
A feature-driven framework for phishing Uniform Resource Locator (URL) detection is introduced, emphasizing the design and evaluation of enhanced feature representations and highlighting that performance gains are primarily driven by feature design rather than model complexity.
Deniz Kaya, Murat Osmanoğlu· PeerJ Computer Science· 0 citations
The Adversarial-Resilient Lightweight Random Forest (AR-LRF) model is proposed, combining controlled ensemble complexity with simulated adversarial perturbations applied during training to mitigate adversarial vulnerabilities.
A. Chaudhuri, M. B· Scientific Reports· 0 citations
In the technology era, Phishing has continued to be a great challenge within the cybersecurity and web security landscape. This involves exploiting human trust on any online services and subtle technical flaws. This is to gather credentials, financial data, and sensitive information across diverse online platforms and various users. Traditional defenses like static blacklists, signature-based filters and simple detection rules are limited by slow update cycles and an inability to capture subtle syntactic and behavioral cues. To address these shortcomings, we propose a hybrid detection framework that fuses classical supervised machine-learning classifiers (e.g., Logistic Regression, SVM, Random Forest, XGBoost) with sequence-aware deep learning (LSTM) to jointly model lexical, structural, syntactic, and behavioral features extracted from URLs and webpage metadata. This combined approach leverages the interpretability and stability of ML models alongside the pattern-learning strength of LSTMs to detect both known and zero-day phishing attempts, produce calibrated confidence scores and deliver comprehensive reports via a real-time web interface resulting in a robust, transparent, and operationally useful solution for strengthening web security.
M. Yaswanth, Pathan Basheer Khan, Dhulipalla Naga Harish et al.· 2026 7th International Confe...· 0 citations
An intelligent phishing detection framework that integrates machine learning, deep learning-enabled feature selection, multi-source phishing feature extraction, hybrid classification, and real-time risk assessment is proposed that can achieve higher detection accuracy, precision, recall, F1-score, and lower response latency than blacklist-based, conventional machine learning, and standalone deep learning approaches.
P. Paul, Bharath Bhushan, Bandameedi Sai Charan et al.· American Journal of AI Cyber...· 0 citations
The proposed LLM-based framework offers a promising approach for improving phishing detection and strengthening modern cybersecurity defenses and suggests that transformer-based models can effectively identify deceptive domain structures, abnormal URL patterns, and obfuscation techniques.
L. Eliyan, M. Alshraideh, Bayan Alfayoumi· Journal of integrated scienc...· 0 citations