White-Box and Transfer-Based Black-Box Adversarial Attacks on Neural Network Models for Credit Card Fraud Detection
Abstract
Machine learning models deployed for credit card fraud detection operate in adversarial, security-critical settings, and their robustness against evasion attacks directly affects financial and operational risk. However, despite extensive work on models for credit card fraud detection, comparatively fewer studies have examined their adversarial vulnerability. This study proposes a targeted Projected Gradient Descent-based attack on feedforward neural network models for credit card fraud detection. The attack is evaluated in white-box and black-box settings, illustrating the extent to which adversarial vectors generated on a surrogate attack model may be transferred to a set of victim models. The experimental settings are designed to ensure a fair, mixed-feature attack, including introduction of perturbation bound constraints, preservation of the class imbalance, specification of attack success criteria, selection of victim models, and avoiding bias. The obtained adversarial transaction vectors are evaluated in terms of their effect on model predictions and their mathematical proximity to the corresponding original fraudulent vectors. Finally, the evaluation results and limitations of the scope of the study are discussed.