Skip to content
Open access

Comparing Machine Learning and Soft-Voting Ensemble Models for Credit Card Fraud Detection in Imbalanced Transaction Data

2026 · American Journal of Student Research · 0 citations

Abstract

Credit card fraud detection is a highly imbalanced classification problem in which headline accuracy can be misleading. In the held-out Kaggle test file used in this study, only 2,145 of 555,719 transactions (0.386%) were fraudulent; a classifier that labeled every transaction legitimate would therefore achieve approximately 99.61% accuracy while detecting no fraud. The analysis evaluated four configured machine-learning pipelines - Decision Tree, Random Forest, XGBoost, and a Multilayer Perceptron (MLP) - together with equal and validation-AP-weighted soft-voting ensembles. Average Precision (AP) was the primary ranking metric, along with precision, recall, F1 score, PR AUC, false positives (FP), and false negatives (FN) used to characterize operating tradeoffs. The selected XGBoost configuration produced the highest observed test AP (0.9084) and F1 score (0.8534) on this synthetic dataset. The soft-voting ensembles reduced some false positives but did not exceed XGBoost in AP. These findings apply to the specific configured pipelines, feature engineering choices, and recurring synthetic customer/ merchant population studied here; they do not establish that XGBoost is generally the best fraud-detection method. The study also identifies limitations related to repeated validation use, unequal imbalance treatments, synthetic-data artifacts, probability calibration, and the absence of uncertainty estimates or entity-disjoint evaluation.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.