Fraud robustness evaluation should report predictive degradation and attack feasibility jointly, and should incorporate domain constraints into attack generation rather than treating them as post-processing checks.
Abstract
Machine learning models are widely used in financial fraud and credit-risk detection, yet their adversarial robustness remains difficult to evaluate because financial tabular data involve domain-specific constraints, severe class imbalance, and asymmetric attacker capability. We argue that, in this setting, robustness is not only an attribute of the model, but also an attribute of the evaluation protocol. Different ways of enforcing constraints and capability can lead to substantially different robustness conclusions. This paper presents FraudBench, a protocol-sensitive benchmark for adversarial robustness evaluation in financial fraud and credit-risk detection. Rather than treating domain constraints as post-hoc validity checks, FraudBench evaluates the same dataset--model--attack--defence setting under three matched protocols: unconstrained attacks, post-hoc feasibility filtering, and deployment-aware constraint-integrated attacks. FraudBench covers four public financial datasets, and evaluates neural, tree-based, and ensemble models using three attack settings. Our results show that robustness conclusions are highly protocol-sensitive. On Lending Club Loan Data under the white-box setting, post-hoc filtering leaves only 3.7 feasible-flipped examples on average, whereas in-attack projection with attacker mutability masking produces 2,832.3 feasible-flipped examples under the same perturbation budget. The results on IEEE-CIS further show that feasibility and attacker capability are separate axes, while black-box evaluation shows that protocol choice can alter model-family rankings. These findings suggest that fraud robustness evaluation should report predictive degradation and attack feasibility jointly, and should incorporate domain constraints into attack generation rather than treating them as post-processing checks.
Machine learning models deployed for credit card fraud detection operate in adversarial, security-critical settings, and their robustness against evasion attacks directly affects financial and operational risk. However, despite extensive work on models for credit card fraud detection, comparatively fewer studies have e...
A. Miljković, Milan Gnjatović, Marijana Joksimović et al.· Electronics· 0 citations
This paper proposes a unified dual-defense framework that jointly integrates adversarial training with a denoising autoencoder (DAE)-based filtering mechanism, specifically designed for imbalanced tabular financial data under adversarial conditions, and explicitly targets adversarial robustness in financial fraud detec...
M. Javeed, Jannatul Maua, M. Mridha et al.· Computers, Materials & C...· 0 citations
Extensive evaluations on seven real-world financial data sets demonstrate that ATDIL outperforms state-of-the-art imbalanced learning methods across multiple metrics while exhibiting superior resilience under adversarial conditions, offering a robust and practical framework for enhancing financial fraud detection syste...
Yu-Hang Tian, Jin Xiao, Le-An Yu et al.· INFORMS journal on computing· 0 citations
Machine learning-based credit scoring is increasingly central to Peer-to-Peer (P2P) lending, yet its resilience to adversarial manipulation, where applicants strategically alter self-reported inputs to secure favourable decisions, remains poorly understood. Most adversarial-robustness evidence comes from image and text...
Gijs A. F. Niewzwaag, Marijn G. S. Veth, Manuele Massei et al.· 0 citations
Machine learning-based network intrusion detection systems (ML-based NIDS) are vulnerable to adversarial evasion, where malicious samples are perturbed to evade detection and be misclassified as benign. Despite growing research on adversarial attacks and defenses for ML-based NIDS, comparative evaluations of multiple a...
Transaction fraud detection remains difficult because fraudulent events are highly imbalanced, temporally drifting, cost asymmetric, and often supported by limited auditability in model decisions. Existing tabular, tree-based, and deep fraud detectors can achieve strong predictive performance, but they often provide we...
Mohammad Sazzad Hossain, Samia Akter, Kaniz Sultana Chy et al.· Discover Data· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.