Skip to content

Author

Kimlong Ngin

3 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Open access Jul 2026

KVerifyID: A Hybrid Multimodal Approach for Khmer Online Writer Verification

Online writer verification with dynamic handwriting signals is still difficult, and it has been especially under-studied forcomplex Southeast Asian scripts like Khmer. This work tackles online, text-independent, word-level Khmer writer verification as apairwise decision problem: given two handwritten word samples, decide whether they were written by the same person. We introduceKVerifyID, a hybrid dual-stream Siamese network that learns from both (i) a grayscale image rendering of each word and (ii) itspen-trajectory sequence (𝑥, 𝑦, 𝑝)with explicit pen-state encoding. The resulting modality embeddings are fused into a compact 128-dimensional writer representation, and verification is performed via cosine similarity, using thresholds selected on validation andthen fixed for testing. On a Khmer online handwriting dataset collected from 298 writers (4,878 word instances) with strict writer-disjoint splits, the model generalizes strongly, achieving 99.50% training accuracy and 99.74% test accuracy, with a low verificationerror of 0.32% validation EER (equal error rate). At the validation equal-error operating point, the errors are FAR (false acceptancerate) = 0.30% and FRR (false rejection rate) = 0.19%. Overall, the results show that jointly leveraging spatial word appearance andonline stroke dynamics enables robust Khmer writer verification, making it promising for digital authentication and forensicscreening.

Kimlong Ngin, Dona Valy, Sokkhey Phauk et al. · 0 citations
Open access Jul 2026

Recognizing Khmer Handwritten Digits with the Power of Sequential RNNs

Recognizing handwritten digits is a fundamental aspect of optical character recognition (OCR), with broad applicationsin areas such as digital archiving, data entry, and assistive technologies. Khmer digits present distinctive challenges due to theirintricate shapes and high variability in individual handwriting styles, making the development of accurate recognition systemsparticularly demanding. Unlike conventional approaches that primarily rely on image-based inputs, this study adopts a sequentialframework using coordinate data, which captures the temporal dynamics of pen strokes and preserves the natural writing sequence.This aspect is especially critical for Khmer digits, where stroke order follows well-defined structural rules. Three recurrent neuralnetwork architectures such as Long Short-Term Memory (LSTM), Bidirectional LSTM (Bi-LSTM), and Gated Recurrent Unit (GRU)were evaluated, with data augmentation applied to improve robustness. A custom dataset of 1,583 handwritten sequences wasexpanded to 12,337 samples through rotation-based augmentation, partitioned into 60% training, 20% validation, and 20% testing.The experimental evaluation reported test accuracies of 95.27% for LSTM, 94.95% for Bi-LSTM, and 95.58% for GRU, with GRUshowing the most promising results. These outcomes validate the effectiveness of sequential modeling for Khmer digit recognitionand emphasize the value of RNN-based methods for complex handwriting systems.

Kimlong Ngin, Dona Valy, Kimhor Phoeurn et al. · 0 citations
Open access Jul 2026

Analysis on Machine Learning Models for Imbalanced Data Problem in Payment Fraud Detection

Payment card fraud losses worldwide reached $33.83 billion in 2023 (Nilson Report, 2025). Alarmingly, Deputy PrimeMinister and Minister of Interior Sar Sokha (2024) stated that in the first semester of 2024, Cambodians lost nearly $40 million toonline and digital fraud. To prevent these significant financial losses, it's crucial to identify the predictive models that can moreaccurately detect the anomalies in transactions. This study investigates which predictive models work best in predicting the anomaliesin the payment transaction. The dataset contains types of online transactions, the amount of the transactions, names of the senderand receiver, the sender’s balance of account balance before and after the transaction, and the receiver’s account balance beforeand after receiving money. Autoencoder, LightGBM, Neural Network, Logistic Regression, Random Forest, and CatBoost were builtas prediction models within this research. Each of these algorithms is employed to create prediction models, which are meticulouslyfine-tuned to yield the most accurate prediction of the fraud payment in the system involving optimizing hyperparameters and selectingthe best features to enhance the models’ prediction power. Common performance metrics such as Precision, Recall, F1-Score, andAUC-ROC were used to test each model's performance. Extensive experimentation data shows the best performance of CatBoost(AUC-ROC of 0.895 for Oversampled and 0.999 for other kinds of datasets) and Random Forest model (AUC-ROC of 0.99), whichconsistently outperforms other machine learning methods in accurately predicting payment fraud, whether training with animbalanced or balanced dataset. These results highlight the model's ability to handle complex, non-linear relationships within thedata and its effectiveness in generalizing across different scenarios and conditions. In conclusion, the findings of this study will bebeneficial mostly to the banking system as they can apply the models we found in their system to prevent any fraudulent activities.Spotting the fraudulent activities in the early stage plays an immense role in preventing the loss.

Siuphing Seun, Sokkhey Phauk, Kimlong Ngin et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.