A Stacking Ensemble-Based Framework for Reliable Phishing Detection in Cybersecurity Risk Mitigation
Abstract
The phishing attack threat has substantially expanded across various digital communication channels and malicious web links, making it a world-wide cyber security phenomenon. Traditional solutions for phishing detection like URL blacklisting do not fight back against elusive zero-day attacks and seamlessly disguised methods. Therefore, this effort intends to take advantage of these vulnerabilities through the design and implementation of a multi-model machine learning solution using stacking ensemble architecture. This study will employ five dissimilar, trusted base learners like Random Forest, Decision Tree, Support Vector Machine (SVM), K-Nearest Neighbors (KNN), Logistic Regression and the Logistic Regression will be the second-level meta-classifier for union and diversity trusted enough in their predictive relevance. Empirical research findings declare how the final developed stacking ensemble model compared favorably and received a greater accuracy level than any single model. Both an accuracy level raised by the implementation of such a model against unseen test data 97.01% compared to F1 score of 97.43% and statistically relative determinations most important was to reduce the False Negative Rate (FNR) to merely 2.19% to minimize the chances that malicious sites will slip through support this. This study also assessed a relative cost-benefit analysis for efficiency with compared security and note that the only caveat within computational costs is, compared to a single decision tree model, the stacker ensemble does have a higher computational cost; however, if anyone using this solution is in a situation where failure to detect those malicious sites is critical, the substantial enhancement in reliability makes it worthwhile.