Skip to content
Open access

Credit Card Fraud Detection under Extreme Class Imbalance: A Comparison of KNN and Logistic Regression

Aug 2026 · UNP Journal of Statistics and Data Science · 0 citations · 12 references

TL;DR

The main finding is that under severe imbalance ROC-AUC can be misleading and PR-AUC is more informative; KNN baseline is a balanced detector without tuning, threshold-tuned LR baseline gives the best single operating point, and LR+SMOTE suits cases where recall is the priority.

Abstract

Credit card fraud is a serious threat in the digital financial ecosystem and is characterised by extreme class imbalance, with fraudulent transactions typically below 1%. This study compares two standard classification algorithms, K-Nearest Neighbor (KNN) and Logistic Regression (LR), for detecting fraudulent transactions on the Sparkov dataset (1.85 million transactions; a stratified subsample of 100,000 rows; 0.52% fraud rate), and analyses the effect of the Synthetic Minority Over-sampling Technique (SMOTE). Preprocessing includes temporal feature engineering, haversine distance, leak-free per-card behavioural features, one-hot and label encoding, and z-score standardisation. Models are evaluated on a stratified 80:20 split using the confusion matrix, accuracy, precision, recall, F1-score, ROC-AUC, and PR-AUC, complemented by decision-threshold tuning, confidence intervals over five repetitions, and the McNemar test. No single model dominates across all metrics. At the default threshold, KNN baseline achieves the highest F1 (0.407) and precision (0.540), while LR baseline achieves the highest PR-AUC (0.246); LR+SMOTE leads on recall (0.712) and ROC-AUC (0.861) but with very low precision (0.025). Threshold tuning lets LR baseline reach the best F1 (0.422 at a 0.071 cut-off). McNemar shows the KNN–LR difference is not significant at baseline (p = 0.282) but significant under SMOTE (p < 0.001). The main finding is that under severe imbalance ROC-AUC can be misleading and PR-AUC is more informative; KNN baseline is a balanced detector without tuning, threshold-tuned LR baseline gives the best single operating point, and LR+SMOTE suits cases where recall is the priority.

Read PDF

Similar papers

Open access Aug 2026

An Enhancing Credit Card Fraud Detection through Data Preprocessing and SMOTE-Based Class Balancing: A Comparative Evaluation of Machine Learning Models

This study compares the performance of several supervised machine learning models for fraud detection, using a unified data preprocessing pipeline, and found that ensemble learning methods generally outperform single classifiers in both accuracy and minority-class recognition.

Nafiu Yahuza, Ahmad Baita Garko, Abubakar Atiku Muslim et al. · 0 citations
Open access Sep 2026

Credit Card Fraud Detection Automation System Using Logistic Regression and XG Boost

An automated machine learning framework for credit card fraud detection that addresses the severe class imbalance inherent in fraud datasets through Random Under-Sampling and Synthetic Minority Over-sampling Technique (SMOTE).

Harshwardhansinh K. Chauhan, Rocky Upadhyay Upadhyay · 0 citations
Review Open access Sep 2026

Credit Card Fraud Detection Using Machine Learning and Risk-Based Alert Strategies

In data environments with severe class imbalances, a machine learning model can be designed as an effective triage tool and its actual value is not only to predict fraud, but also to establish a low-risk release, medium-risk verification, and high-risk manual review of the decision-making process for institutions and t...

Zi-Yue Meng · 0 citations
Open access 2026

Comparing Machine Learning and Soft-Voting Ensemble Models for Credit Card Fraud Detection in Imbalanced Transaction Data

Credit card fraud detection is a highly imbalanced classification problem in which headline accuracy can be misleading. In the held-out Kaggle test file used in this study, only 2,145 of 555,719 transactions (0.386%) were fraudulent; a classifier that labeled every transaction legitimate would therefore achieve approxi...

Christian Barravecchio · 0 citations
Open access Aug 2026

Neural Network-based Model for Detecting Credit Card Fraud: A Comparative Study of Oversampling Techniques and Feature Selection

Credit card fraud detection is complicated by the severe class imbalance typical of transaction data, because fraudulent cases represent only a small proportion of observations. This study develops a neural-network-based model for classifying transactions as legitimate or fraudulent and compares combinations of two ove...

Kalala Kanyinda Norbert, Mukala Patrick, Kafunda Katalayi Pierre · 0 citations
#explainable ai Open access Sep 2026

Explainable Fraud Detection AI System in Financial Sector

The study shows that ensemble models on the original feature space provide highly accurate and stable fraud detection on this dataset and SHAP analysis reveals that source and destination balances, transaction amount and type are the most influential features.

Merit Chinonso Opara · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.