Skip to content
Open access

Machine learning for fraud detection in public sector financial systems

Sep 2026 · Finance & Accounting Research Journal · 0 citations

TL;DR

This study examines whether machine learning can improve detection of fraudulent transactions in public sector financial systems and compares model families under conditions resembling real government data, including severe class imbalance and scarce fraud labels.

Abstract

Purpose: Public sector financial systems process large volumes of payments, procurement transactions and grants and remain exposed to fraud, corruption and error. Traditional audit-based controls detect only a small share of irregular transactions and usually do so late. This study examines whether machine learning can improve detection of fraudulent transactions in public sector financial systems and compares model families under conditions resembling real government data, including severe class imbalance and scarce fraud labels. Design/methodology/approach: The study builds a synthetic public-sector transaction dataset of 60,000 records that reproduces documented properties of government payment data, including low fraud prevalence, vendor concentration and period-end clustering. Four models are trained and compared: logistic regression, random forest, XGBoost and isolation forest. Models are evaluated using precision, recall, F1 score, area under the receiver operating characteristic curve and area under the precision-recall curve, since accuracy alone misleads under imbalance. Findings: XGBoost achieved the strongest balance of precision and recall (F1 = 0.762, AUC-ROC = 0.992), followed by random forest (F1 = 0.748). Logistic regression achieved high recall but low precision, producing many false alerts. Isolation forest, which is an unsupervised method, performed reasonably well without labelled fraud, which matters since most agencies lack reliable labels. Amount, vendor age, single-bidder procurement and prior irregularity counts were the strongest predictors. Originality/value: This study contributes a reproducible, coded evaluation framework for public sector fraud analytics and gives practitioners evidence-based guidance on model choice under label scarcity and class imbalance, which are conditions that dominate real government settings. Keywords: Fraud Detection, Machine Learning, Public Sector, Financial Systems, Public Procurement, Anomaly Detection, Government Auditing, XGBoost.

Read PDF

Similar papers

Open access 2026

Intelligent Real-Time Fraud Detection in Financial Institutions

The results show that the stacked ensemble produced a usable prototype-level fraud decision layer by combining supervised and anomaly based evidence, although threshold calibration, explainability, and validation on local institutional data remain necessary before operational deployment.

Nwadike U. S., Emmah V. T., M. D. · 0 citations
Review Open access Sep 2026

Credit Card Fraud Detection Using Machine Learning and Risk-Based Alert Strategies

In data environments with severe class imbalances, a machine learning model can be designed as an effective triage tool and its actual value is not only to predict fraud, but also to establish a low-risk release, medium-risk verification, and high-risk manual review of the decision-making process for institutions and t...

Zi-Yue Meng · 0 citations
#explainable ai Open access Sep 2026

Explainable Fraud Detection AI System in Financial Sector

The study shows that ensemble models on the original feature space provide highly accurate and stable fraud detection on this dataset and SHAP analysis reveals that source and destination balances, transaction amount and type are the most influential features.

Merit Chinonso Opara · 0 citations
Open access 2023

Reducing Financial Fraud Using Machine Learning and CRM Data Models

The fact that banks may find fraud, minimize risks before they happen, and make their clients happy by combining ML and CRM standard data models together in an effective manner is demonstrated.

Satyendra Kumar Vanapalli · 0 citations
Open access Sep 2026

Credit Card Fraud Detection Automation System Using Logistic Regression and XG Boost

An automated machine learning framework for credit card fraud detection that addresses the severe class imbalance inherent in fraud datasets through Random Under-Sampling and Synthetic Minority Over-sampling Technique (SMOTE).

Harshwardhansinh K. Chauhan, Rocky Upadhyay Upadhyay · 0 citations
Open access Sep 2026

CREDIT CARD FRAUD DETECTION USING ENSEMBLE MACHINE LEARNING METHODS

The study results show that ensemble techniques enhance the detection of fraudulent credit card transactions significantly and the proposed framework included data preprocessing and imbalanced data handling using SMOTE produced good results in terms of classification performance and reliable detections.

Moohanad Jawthari, Ihsan Sahib, N. H. Fadhil · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.