Skip to content

Author

Muhammad Hilmy Aziz

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Open access Aug 2026

Comparative Analysis of Gradient Boosting Models for Phishing Website Detection

The growth of digital platforms has unfortunately led to the proliferation of social engineering deceptions, especially in the form of phishing. Existing security techniques that rely on static lists are not effective at identifying freshly registered rogue domains and hence leave internet users extremely vulnerable. To address this limitation, this study develops and evaluates an automated phishing website detection system using advanced gradient boosting algorithms, namely Extreme Gradient Boosting, Light Gradient Boosting Machine, and Categorical Boosting. The methodology utilizes the pre-structured Mendeley Phishing Websites Dataset, where comprehensive hybrid features from both uniform resource locator characteristics and hypertext markup language source code were already extracted. The models are optimized using a randomized search method combined with five-fold stratified cross validation, and predictive performance is systematically evaluated across four distinct data partitioning scenarios to ensure robustness. The experimental evaluations reveal that all models achieve their peak classification performance at the 80:20 training and testing data distribution. Among the evaluated frameworks, Categorical Boosting demonstrates the most superior capability, achieving an overall accuracy of 97.40 percent, a recall of 96.11 percent, and an F1 score of 96.12 percent. The high recall metric proves the practical effectiveness of Categorical Boosting in minimizing false negative predictions, ensuring that zero-day threats do not bypass the system undetected. These findings establish Categorical Boosting's performance as strong benchmark evidence on a single public dataset; however, further validation on newer campaigns, unseen domains, adversarially obfuscated sites, and cross-dataset settings is still required to confirm its generalization capabilities.

Muhammad Hilmy Aziz, Rizky Parlika, Kartini Kartini · 0 citations