Cost-sensitive random forest with adaptive threshold optimization for imbalanced financial fraud detection
Abstract
Financial transaction fraud detection is a challenging binary classification problem characterized by severe class imbalance, rare fraudulent cases, and asymmetric misclassification costs. In real-world banking systems, failing to detect fraudulent transactions can lead to substantial financial losses, whereas false alarms mainly increase operational review costs. Therefore, fraud detection models should not be optimized only for overall accuracy, as accuracy-oriented classifiers may favor the majority non-fraud class and fail to identify minority fraud cases effectively. To address this issue, this study proposes a cost-sensitive ML framework for financial transaction fraud detection under class imbalance and asymmetric financial risk. In this study, a cost-sensitive RF framework with adaptive probability threshold optimization is proposed for binary fraud classification. Misclassification costs are incorporated during model training through class-weighted learning, allowing the classifier to assign greater importance to minority fraud cases. In addition, instead of using the default probability threshold of 0.5, the decision threshold is optimized using validation data to minimize the expected misclassification cost and improve minority-class detection. The proposed framework was evaluated on two publicly available financial fraud datasets with different transaction characteristics: the Credit Card Fraud Detection dataset and the Financial Fraud Prediction dataset. Its performance was compared with conventional ML baselines using fraud-oriented evaluation metrics, including precision, recall, F1-score, false-negative rate, false-positive rate, and expected misclassification cost. The proposed framework achieved strong and consistent performance across both datasets. On the Credit Card Fraud Detection dataset, it achieved a precision of 94.4%, recall of 85.7%, and F1-score of 89.8%. On the Financial Fraud Prediction dataset, it achieved a precision of 93.1%, a recall of 83.8%, and an F1-score of 88.2%. Compared with conventional fixed-threshold ML models, the adaptive threshold strategy reduced false-negative predictions while maintaining high precision. These results indicate that the proposed cost-sensitive framework provides a more balanced decision strategy for fraud detection under highly imbalanced transaction distributions. The findings demonstrate that combining class-weighted learning with adaptive threshold optimization can improve fraud detection performance under asymmetric misclassification costs. The proposed framework offers a practical and interpretable decision-support approach for banking fraud detection systems, particularly in settings where reducing undetected fraudulent transactions is more important than maximizing overall accuracy alone.