Skip to content
Open access

Detection of Hate and Dehumanizing Speech in Pashto Text Using Machine Learning Algorithms

Sep 2026 · Journal of nature and science · 0 citations · 26 references

Abstract

Detecting harmful language in online platforms is essential for creating safer and more inclusive social media environments. However, research on harmful language detection has largely focused on high-resource languages, while low-resource languages such as Pashto remain comparatively underexplored. This study compares hate speech and dehumanizing language in Pashto social media comments. To address this gap, we developed a dataset containing 1,590 Pashto comments collected from social media and manually categorized them into four classes: Hate, Dehumanization, Both, and Neither. Three traditional machine learning algorithms, Logistic Regression, Naive Bayes, and Linear Support Vector Machine (SVM), were evaluated using Term Frequency-Inverse Document Frequency (TF-IDF) features. Among the evaluated models, Logistic Regression achieved the best overall performance, reaching an accuracy of 71.63%. The results indicate that distinguishing between hate speech and dehumanizing language remains challenging, particularly when both forms occur within the same comment. The model showed a tendency to confuse the Hate and Dehumanization categories, with the Both class presenting additional classification difficulties. Furthermore, the experiments demonstrate that appropriate text normalization and selecting an effective feature size, approximately 3,500 features in this study, can improve classification performance.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.