Optimizing Fraud Detection in Financial Transaction Systems with Machine Learning
Abstract
Financial fraud remains a critical challenge in the banking and finance industry, leading to annual losses exceeding $32 billion globally. This paper introduces an approach to enhancing fraud detection in financial transaction systems using machine learning models. Integrating systems engineering principles with advanced analytics, we developed a scalable, real-time fraud detection framework for a mid-sized financial institution.By analyzing existing transaction workflows, we identified bottlenecks and vulnerabilities. Restructuring the data architecture with normalized relational schemas and dimensional modeling in SQL improved data retrieval speeds by 30%, enhancing data management efficiency.Automated ETL pipelines via Azure Data Factory reduced data processing time by 45% and minimized manual errors. Machine learning algorithms—including Random Forest and XGBoost—were trained on over 15 million transactions spanning five years. Engineered features such as transaction amount, frequency, geolocation, and merchant category enhanced model performance.The models achieved a fraud detection accuracy of 96% and reduced false positives by 20% compared to the prior system. Real-time processing with Databricks and Apache Spark enabled handling up to 12,000 transactions per second with latency under 150 milliseconds.Implementing this optimized system resulted in a 25% reduction in fraud-related losses and a 22% increase in transaction processing efficiency. This project demonstrates the effective application of industrial engineering methodologies to improve complex financial systems through advanced analytics and cloud technologies, ultimately enhancing operational performance and security.