Skip to content

Author

Pakrigna Long

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Open access Jul 2026

Analysis on Machine Learning Models for Imbalanced Data Problem in Payment Fraud Detection

Payment card fraud losses worldwide reached $33.83 billion in 2023 (Nilson Report, 2025). Alarmingly, Deputy PrimeMinister and Minister of Interior Sar Sokha (2024) stated that in the first semester of 2024, Cambodians lost nearly $40 million toonline and digital fraud. To prevent these significant financial losses, it's crucial to identify the predictive models that can moreaccurately detect the anomalies in transactions. This study investigates which predictive models work best in predicting the anomaliesin the payment transaction. The dataset contains types of online transactions, the amount of the transactions, names of the senderand receiver, the sender’s balance of account balance before and after the transaction, and the receiver’s account balance beforeand after receiving money. Autoencoder, LightGBM, Neural Network, Logistic Regression, Random Forest, and CatBoost were builtas prediction models within this research. Each of these algorithms is employed to create prediction models, which are meticulouslyfine-tuned to yield the most accurate prediction of the fraud payment in the system involving optimizing hyperparameters and selectingthe best features to enhance the models’ prediction power. Common performance metrics such as Precision, Recall, F1-Score, andAUC-ROC were used to test each model's performance. Extensive experimentation data shows the best performance of CatBoost(AUC-ROC of 0.895 for Oversampled and 0.999 for other kinds of datasets) and Random Forest model (AUC-ROC of 0.99), whichconsistently outperforms other machine learning methods in accurately predicting payment fraud, whether training with animbalanced or balanced dataset. These results highlight the model's ability to handle complex, non-linear relationships within thedata and its effectiveness in generalizing across different scenarios and conditions. In conclusion, the findings of this study will bebeneficial mostly to the banking system as they can apply the models we found in their system to prevent any fraudulent activities.Spotting the fraudulent activities in the early stage plays an immense role in preventing the loss.

Siuphing Seun, Sokkhey Phauk, Kimlong Ngin et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.