Skip to content
Open access

Adversarial Robustness of Machine Learning-based Fraud Detection Systems: An Empirical Evaluation of Attack Impact and Mitigation in Fintech Environments

Jul 2026 · Journal of Engineering Research and Reports · Vol 28, pp. 261-279 · 0 citations

TL;DR

The findings support the adoption of adversarial training for gradient-based fraud detection models and suggest that robustness claims should be accompanied by disclosure of the attack methodology used to establish them.

Abstract

This study evaluated the adversarial robustness of machine learning-based fraud detection systems by comparing classifier vulnerability profiles and assessing adversarial training as a mitigation strategy. Using the IEEE-CIS Fraud Detection dataset, comprising 590,540 transactions with a fraud incidence of 3.5%, four classifiers—logistic regression, random forest, gradient boosting, and a feed-forward neural network—were trained under identical preprocessing and class-weighting conditions and then subjected to Fast Gradient Sign Method and Projected Gradient Descent attacks at a perturbation budget of 0.02. Adversarial examples were constructed directly using closed-form and backpropagated gradients for the differentiable classifiers and using a logistic regression surrogate for the non-differentiable ensembles, before adversarial training was applied as a post-attack mitigation stage. Logistic regression proved the most adversarially vulnerable architecture, sustaining a 31.12-percentage-point recall loss under Projected Gradient Descent, while adversarial training subsequently restored its recall from 0.39 to 0.999 at an accuracy cost of 0.10 percentage points. Random forest and gradient boosting were not degraded by the surrogate-based attack, indicating that comparative robustness claims for tree-based ensembles require attack methods suited to their non-differentiable structure rather than transfer-based evaluation alone. Within the scope of this single-dataset evaluation, the findings support the adoption of adversarial training for gradient-based fraud detection models and suggest that robustness claims should be accompanied by disclosure of the attack methodology used to establish them.

Read PDF

Similar papers

Open access Sep 2026

White-Box and Transfer-Based Black-Box Adversarial Attacks on Neural Network Models for Credit Card Fraud Detection

Machine learning models deployed for credit card fraud detection operate in adversarial, security-critical settings, and their robustness against evasion attacks directly affects financial and operational risk. However, despite extensive work on models for credit card fraud detection, comparatively fewer studies have e...

A. Miljković, Milan Gnjatović, Marijana Joksimović et al. · 0 citations
Open access 2026

A Two-Stage Adversarial Defense Architecture for Robust Fraud Detection on Imbalanced Financial Data

This paper proposes a unified dual-defense framework that jointly integrates adversarial training with a denoising autoencoder (DAE)-based filtering mechanism, specifically designed for imbalanced tabular financial data under adversarial conditions, and explicitly targets adversarial robustness in financial fraud detec...

M. Javeed, Jannatul Maua, M. Mridha et al. · 0 citations
Preprint Sep 2026

Robustness Evaluation and Detection of Transferable Adversarial Attacks in ML-Based NIDS

Machine learning-based network intrusion detection systems (ML-based NIDS) are vulnerable to adversarial evasion, where malicious samples are perturbed to evade detection and be misclassified as benign. Despite growing research on adversarial attacks and defenses for ML-based NIDS, comparative evaluations of multiple a...

Huda Ali Alatawi · 0 citations
Open access Sep 2026

Security Evaluation of Classical Machine Learning Models Under Poisoning, Evasion, and Model Extraction Attacks

Classical machine learning (ML) models, including Logistic Regression (LR), Support Vector Machines (SVM), Random Forests (RF), and XGBoost, remain widely used in practical applications because of their efficiency, interpretability, and relatively low computational costs. However, their security properties against diff...

Hannelore Sebestyen, Elisa Valentina Moisi, S. Coman et al. · 0 citations
Sep 2026

Adversarial Training–Based Deep Imbalanced Learning

Extensive evaluations on seven real-world financial data sets demonstrate that ATDIL outperforms state-of-the-art imbalanced learning methods across multiple metrics while exhibiting superior resilience under adversarial conditions, offering a robust and practical framework for enhancing financial fraud detection syste...

Yu-Hang Tian, Jin Xiao, Le-An Yu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.