Skip to content
Open access

Performance Evaluation of Classical Machine Learning Models for Insurance Fraud Detection Under Severe Class Imbalance

Aug 2026 · International Journal For Multidisciplinary Research · 0 citations · 18 references

TL;DR

Experimental results demonstrate that ensemble-based methods, particularly Random Forest, achieve superior performance in identifying minority class fraud cases while maintaining stable overall classification accuracy in fraud detection scenarios.

Abstract

Insurance claim fraud continues to pose substantial financial and operational challenges for insurance companies worldwide, particularly in auto insurance, where fraudulent cases form a small but highly impactful portion of overall claims. One of the significant technical difficulties in detecting such fraud is the severe class imbalance in real-world insurance datasets, where legitimate claims vastly outnumber fraudulent ones. This study presents a systematic performance evaluation of classical machine learning (ML) classification models for insurance fraud detection under conditions of extreme class imbalance. The proposed framework focuses on widely used supervised learning algorithms, including Support Vector Machine (SVM), K-Nearest Neighbor (KNN), and Random Forest (RF), with an emphasis on understanding their behavior when trained on imbalanced data. To mitigate class imbalance bias, the Synthetic Minority Oversampling Technique (SMOTE) is applied to the dataset before model training. Model performance is evaluated using multiple metrics such as accuracy, precision, recall, and F1-score, which provide a more reliable assessment than accuracy alone in fraud detection scenarios. Experimental results demonstrate that ensemble-based methods, particularly Random Forest, achieve superior performance in identifying minority class fraud cases while maintaining stable overall classification accuracy. This research provides practical insights into selecting suitable classical ML models for insurance fraud detection, supporting the development of reliable decision support systems for insurance providers operating with imbalanced data.

Read PDF

Similar papers

Open access Oct 2026

EXPLAINABLE MACHINE LEARNING FOR INSURANCE CLAIM FRAUD DETECTION: A COMPARATIVE PYTHON-BASED STUDY

Today insurance fraud is a big problem for insurance companies since it is challenging to detect potential fraud cases within a large amount of claims data. Traditional rule-based models can detect already known fraud cases, but they often fail to detect new ones and can flag too many false positives. In this paper, se...

Ermira Memeti · 0 citations
Open access Sep 2026

CREDIT CARD FRAUD DETECTION USING ENSEMBLE MACHINE LEARNING METHODS

The study results show that ensemble techniques enhance the detection of fraudulent credit card transactions significantly and the proposed framework included data preprocessing and imbalanced data handling using SMOTE produced good results in terms of classification performance and reliable detections.

Moohanad Jawthari, Ihsan Sahib, N. H. Fadhil · 0 citations
Review Open access Oct 2026

Cost-sensitive random forest with adaptive threshold optimization for imbalanced financial fraud detection

Financial transaction fraud detection is a challenging binary classification problem characterized by severe class imbalance, rare fraudulent cases, and asymmetric misclassification costs. In real-world banking systems, failing to detect fraudulent transactions can lead to substantial financial losses, whereas fals...

Jafar Ali, Rahman Shafique, Daniel Gavilanes et al. · 0 citations
Open access Aug 2026

INSURANCE FRAUD DETECTION USING MACHINE LEARNING ON CLASSIMBALANCED DATASETS WITH MISSING VALUES

The findings show that correcting class imbalance is essential to enhancing model performance, and handling missing data also helps to produce predictions that are more trustworthy, and the AdaBoost Classifier greatly outperforms current methods.

Uzma Fatima, Lubna Nausheen, Sadaf Jahan · 0 citations
Open access Aug 2026

An Enhancing Credit Card Fraud Detection through Data Preprocessing and SMOTE-Based Class Balancing: A Comparative Evaluation of Machine Learning Models

This study compares the performance of several supervised machine learning models for fraud detection, using a unified data preprocessing pipeline, and found that ensemble learning methods generally outperform single classifiers in both accuracy and minority-class recognition.

Nafiu Yahuza, Ahmad Baita Garko, Abubakar Atiku Muslim et al. · 0 citations
Open access Aug 2026

Credit Card Fraud Detection under Extreme Class Imbalance: A Comparison of KNN and Logistic Regression

The main finding is that under severe imbalance ROC-AUC can be misleading and PR-AUC is more informative; KNN baseline is a balanced detector without tuning, threshold-tuned LR baseline gives the best single operating point, and LR+SMOTE suits cases where recall is the priority.

Felicia Sword, Christopher Andreas · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.