Aug 2026· International Journal For Multidisciplinary Research· 0 citations· 18 references
TL;DR
Experimental results demonstrate that ensemble-based methods, particularly Random Forest, achieve superior performance in identifying minority class fraud cases while maintaining stable overall classification accuracy in fraud detection scenarios.
Abstract
Insurance claim fraud continues to pose substantial financial and operational challenges for insurance companies worldwide, particularly in auto insurance, where fraudulent cases form a small but highly impactful portion of overall claims. One of the significant technical difficulties in detecting such fraud is the severe class imbalance in real-world insurance datasets, where legitimate claims vastly outnumber fraudulent ones. This study presents a systematic performance evaluation of classical machine learning (ML) classification models for insurance fraud detection under conditions of extreme class imbalance. The proposed framework focuses on widely used supervised learning algorithms, including Support Vector Machine (SVM), K-Nearest Neighbor (KNN), and Random Forest (RF), with an emphasis on understanding their behavior when trained on imbalanced data. To mitigate class imbalance bias, the Synthetic Minority Oversampling Technique (SMOTE) is applied to the dataset before model training. Model performance is evaluated using multiple metrics such as accuracy, precision, recall, and F1-score, which provide a more reliable assessment than accuracy alone in fraud detection scenarios. Experimental results demonstrate that ensemble-based methods, particularly Random Forest, achieve superior performance in identifying minority class fraud cases while maintaining stable overall classification accuracy. This research provides practical insights into selecting suitable classical ML models for insurance fraud detection, supporting the development of reliable decision support systems for insurance providers operating with imbalanced data.
Today insurance fraud is a big problem for insurance companies since it is challenging to detect potential fraud cases within a large amount of claims data. Traditional rule-based models can detect already known fraud cases, but they often fail to detect new ones and can flag too many false positives. In this paper, se...
Ermira Memeti· INTERNATIONAL CONGRESS FROM...· 0 citations
The study results show that ensemble techniques enhance the detection of fraudulent credit card transactions significantly and the proposed framework included data preprocessing and imbalanced data handling using SMOTE produced good results in terms of classification performance and reliable detections.
Moohanad Jawthari, Ihsan Sahib, N. H. Fadhil· Journal of Digital Security...· 0 citations
Financial transaction fraud detection is a challenging binary classification problem characterized by severe class imbalance, rare fraudulent cases, and asymmetric misclassification costs. In real-world banking systems, failing to detect fraudulent transactions can lead to substantial financial losses, whereas fals...
Jafar Ali, Rahman Shafique, Daniel Gavilanes et al.· Frontiers in Artificial Inte...· 0 citations
The findings show that correcting class imbalance is essential to enhancing model performance, and handling missing data also helps to produce predictions that are more trustworthy, and the AdaBoost Classifier greatly outperforms current methods.
Uzma Fatima, Lubna Nausheen, Sadaf Jahan· International Journal of AI...· 0 citations
This study compares the performance of several supervised machine learning models for fraud detection, using a unified data preprocessing pipeline, and found that ensemble learning methods generally outperform single classifiers in both accuracy and minority-class recognition.
Nafiu Yahuza, Ahmad Baita Garko, Abubakar Atiku Muslim et al.· Lead Sci Journal of Manageme...· 0 citations
The main finding is that under severe imbalance ROC-AUC can be misleading and PR-AUC is more informative; KNN baseline is a balanced detector without tuning, threshold-tuned LR baseline gives the best single operating point, and LR+SMOTE suits cases where recall is the priority.
Felicia Sword, Christopher Andreas· UNP Journal of Statistics an...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.