Skip to content
Open access

Temporal Evaluation of Gradient Boosting and Autoencoder Models for Credit Card Fraud Detection

Sep 2026 · International Journal For Multidisciplinary Research · 0 citations · 21 references

Abstract

Payment card fraud losses exceeded US$33 billion worldwide in 2022, yet fraud detection models are often summarized by accuracy values that conceal how much fraud they miss. This study compares a Naive Bayes baseline, three gradient boosting libraries (LightGBM, XGBoost and CatBoost), an autoencoder anomaly detector, and a hybrid model that supplies the autoencoder's reconstruction error to LightGBM, using the IEEE-CIS Fraud Detection dataset. The 590,540 labeled transactions were ordered in time and divided into training (60%), validation (20%) and test (20%) periods, and all preprocessing was fitted on the training period only. On the later test period (118,108 transactions, 4,064 of them fraudulent), LightGBM reached an average precision (AP) of 0.501 and a ROC-AUC of 0.889. XGBoost reached an AP of 0.506 when text features were supplied as integer codes but only 0.325 with its native categorical handling, a larger gap than any observed between libraries. CatBoost reached 0.460 at its 800-iteration cap. The autoencoder alone was weaker than Naive Bayes (AP 0.111 versus 0.132), and adding its reconstruction error to LightGBM changed AP by +0.003 (95% day-bootstrap interval −0.002 to 0.009). At thresholds selected on validation data, the strongest models detected 43–46% of test fraud at 53–55% precision, whereas a classifier that flags nothing reached 96.6% accuracy. Every model performed worse on the test period than on validation. These results indicate that categorical encoding, threshold choice and temporal evaluation affect reported fraud detection performance at least as much as the choice among modern boosting libraries.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.