Skip to content
Open access

Arabic Plagiarism Detection Using Word2Vec-Based Semantic Features and Random Forest Classification on the ExAraPlagDet Dataset

Aug 2026 · Informatica · 0 citations · 35 references

TL;DR

The findings underscore the potential of advanced NLP techniques to overcome language-specific challenges, providing a foundation for future research in multilingual plagiarism detection and enhancing the development of tools for other languages facing similar challenges.

Abstract

Plagiarism detection has become a critical challenge in the digital age, particularly for languages with complex structures such as Arabic. Traditional methods relying on string matching and basic lexical analysis, are insufficient for detecting more sophisticated forms of plagiarism like paraphrasing and synonym substitution in Arabic texts. This research addresses this gap by proposing a novel approach that employs that employs word embedding and semantic similarity measures within a machine learning framework. Specifically, we utilize models such as Support Vector Machines (SVM) and neural networks to capture the nuanced semantic relationships between words, enabling more effective detection of subtle semantic similarities in text. Our methodology encompasses the development and evaluation of machine learning models, specifically tailored to the unique characteristics of the Arabic language [1]. We conducted extensive experiments on a diverse dataset of Arabic texts, consisting of over 50,000 documents from various sources, including academic publications, online articles, and literary works, to demonstrate the effectiveness of our approach. Our results demonstrate substantial improvements in both accuracy and robustness, surpassing traditional plagiarism detection techniques. The Random Forest classifier achieved the best performance with precision, recall, and F1-score all reaching 0.96, significantly outperforming Decision Tree, Logistic Regression, and SVM. These results confirm the superiority of the Random Forest approach for the given classification problem. This study contributes to the field by developing a more effective tool for plagiarism detection, which is crucial for maintaining academic integrity and protecting intellectual property in Arabic-speaking communities. The findings underscore the potential of advanced NLP techniques to overcome language-specific challenges, providing a foundation for future research in multilingual plagiarism detection and enhancing the development of tools for other languages facing similar challenges.

Read PDF

Similar papers

Open access Aug 2026

Evaluating lexical feature extraction for plagiarism detection in Arabic documents

This study introduces an external plagiarism detection framework built on an artificial neural network model and a lexical feature extraction framework adapted to the linguistic features of Arabic, verifying its effectiveness for Arabic plagiarism detection.

Marwah Alian, Dana Halabi, H. Alshboul · 0 citations
Jul 2026

A novel semantic–syntactic hybrid plagiarism detection system based on word embeddings and similarity measures

Experimental results indicate that the proposed approach achieves competitive performance compared to existing plagiarism detection systems, and the comparative analysis highlights the strengths and limitations of different word embedding models across datasets.

Malya Singh, Vishal Gupta · 0 citations
Open access Jul 2026

LSTM-Based Classification of Indonesian Regional Song Lyrics by Language

This study successfully proposes a Long Short-Term Memory (LSTM)-based model for automatic classification of Indonesian regional song lyrics by language, demonstrating that LSTM effectively captures sequential linguistic patterns and contextual relationships within regional languages.

Muhammad Rizky, Anandita Priatama, Aviv Yuniar Rahman et al. · 0 citations
Open access Jul 2026

Multilingual AI-Generated Text Detection in Arabic, English, and Turkish Using a Hybrid Transformer–Graph Convolutional Network

A hybrid architecture that combines a Transformer-based DistilBERT model with a Graph Convolutional Network (GCN) that enhances detection by modeling structural relationships within text data is proposed.

Ayca Bostancioglu, Bihter Das, Muzeyyen Bulut Ozek · 0 citations
Open access 2026

Multilingual Plagiarism Detection Using GNNs and Syntax-Semantic Knowledge Graphs

A new hybrid Approach for CLPD is proposed, which combines semantic information from WordNet with the syntactic structure from Universal Dependencies, then these relations are modeled in knowledge graphs for multiple language pairs, demonstrating clear improvements over state-of-the-art baselines.

Chaimaa Bouaine, F. Benabbou, Amine Bouaine et al. · 0 citations
Open access Jun 2026

Language Identification in Transliteration-Based Code-Mixed Text: A Study on Telugu–English Data

This work focuses on word-level Language Identification (LID) for transliterated text in informal Roman transliteration, and relies on character-based TF–IDF features and a set of traditional machine-learning models.

Adarshavathi Jampala, Padmavathi Guddetti, G. Kancharla · 0 citations