Enhancing Fraudulent Transaction Detection Through SMOTE-Based Data Balancing and Machine Learning
Abstract
Financial fraud has emerged as one of the most pressing challenges in the digital economy, causing billions of dollars in losses annually across global banking and e-commerce sectors. The highly imbalanced nature of fraud datasets, where legitimate transactions vastly outnumber fraudulent ones, poses significant obstacles for traditional machine learning classifiers. This research investigates the application of Synthetic Minority Over-sampling Technique (SMOTE) for addressing class imbalance in fraudulent transaction detection systems. We implemented and evaluated multiple machine learning algorithms including Random Forest, Gradient Boosting, Logistic Regression, and Support Vector Machines on credit card transaction datasets before and after SMOTE application. Our experimental results demonstrate that SMOTE-enhanced models achieve substantial improvements in fraud detection performance, with Random Forest showing the highest accuracy of 99.2% and F1-score of 0.89 on balanced data compared to 97.4% accuracy and 0.42 F1-score on imbalanced data. The study reveals that data balancing techniques significantly enhance minority class detection without compromising overall model performance, making them essential for real-world fraud detection applications where missing fraudulent transactions carries severe financial and reputational consequences.