Effectiveness of the Random Forest Classifier Algorithm in Credit Card Fraud Detection in Financial Institutions
Abstract
The increase in the number of credit card users and transactions has caused fraudulent activities to increase significantly. As a result, it has become crucial for financial institutions to develop robust and accurate fraud detection systems to minimise financial losses. One of the popular ways of achieving this is through machine learning. The objective of this research is to develop a model that can detect whether a credit card used for a transaction has been compromised. In this study, a detection system was developed using the random forest classifier. The classifier was used to analyse one million, two hundred and ninety-six thousand, six hundred and seventy-five (1,296,675) synthetic credit card transactions generated with the Sparkov Data Generator and obtained from Kaggle. The dataset contained one million, two hundred and eighty-nine thousand, one hundred and sixty-nine (1,289,169) non-fraudulent transactions and seven thousand five hundred and six (7,506) fraudulent transactions. The data was split with 70% used for training and 30% for testing before being passed into the random forest classifier. The training data was used to fit the model while the testing data was used to evaluate it, and a confusion matrix was plotted. A feature importance plot showed the factors with the highest importance in the dataset. The performance of the algorithm was analysed using accuracy, precision, recall, and F1-score, with fraudulent transactions treated as the positive class. The software used was the Python programming language and the scikit-learn module. The model achieved an overall accuracy of 99.76%, a precision of 91.40%, a recall of 64.20%, and an F1-score of 75.43%, correctly identifying 64.20% of fraudulent transactions and 99.96% of non-fraudulent transactions. The results demonstrate that the random forest classifier is highly precise but that recall remains constrained by severe class imbalance, indicating that resampling or cost-sensitive learning is required before deployment. The model developed contributes to the management of credit card fraud in financial institutions.