Skip to content
Review Open access

Detecting Hate Speech in Hindi Digital Discourse Using Transformer–Long Short-Term Memory Models

Jul 2026 · Journal of Communication, Language and Culture · 0 citations · 29 references

TL;DR

Evaluating advanced language models for analysing hate speech in Hindi social media content illustrates that language-specific computational tools can be used for both platform governance and communication research, provided that the cultural context is considered.

Abstract

Hate speech on digital communication platforms has become a major obstacle to healthy online discourse, especially in multilingual societies such as India, where Hindi is a dominant language in social media interactions. Hostile, offensive, and defamatory speech is linguistically and socio-culturally complex because of colloquial idioms, regional variations, code-mixing, and culturally embedded references. However, effective detection is crucial for creating safer and more inclusive digital communication environments. This study evaluates advanced language models for analysing hate speech in Hindi social media content. A dataset of 21000 Hindi posts from Twitter and public repositories, categorized into five categories, was analysed: hate, offensive, fake, defamation, and non-hostile. General-purpose models were tested against Hindi-specific language models, including Hindi-BERT (based on Bidirectional Encoder Representations from Transformers, or BERT) and MuRIL, to investigate whether performance can be further enhanced by integrating Long Short-Term Memory (LSTM) layers. In tests, the best-performing model, Hindi-BERT (with sequential learning added), correctly identified 94 out of every 100 posts—a considerable improvement over simpler models. For social media platforms, it has real-world consequences: it can automatically identify harmful content to be reviewed by a human, minimize nuisance alerts that waste human moderators' time, and identify defamatory or bogus posts before they gain a lot of traction. The results offer a systematic approach to researching how hostility develops in the Hindi-speaking online community, how linguistic creativity (e.g., slang, sarcasm, code-mixing) can conceal or manifest hostility, and how decisions about content moderation influence public discourse for communication scholars. Overall, this paper illustrates that language-specific computational tools can be used for both platform governance and communication research, provided that the cultural context is considered. Finally, technical methods are combined with communication scholarship to explain and curb harmful speech in the online public sphere.

Read PDF

Similar papers

Open access Jul 2026

Low-Resource Hate Speech Detection in English-Swahili Code-Switched Text Using Fine-Tuning of Pre-trained Language Models

This study explores a low-resource approach to detecting hate speech in English and Swahili code-switched text by fine-tuning pre-trained language models, and shows that fine-tuning modern language models can offer a practical and scalable solution for hate speech detection in multilingual environments.

Kipkebut Andrew, Jepkemei Betty · 0 citations
Review Open access 2026

Culturally Aware Malay–English Code-Mixed Hate Speech Detection: A Systematic Review and Research Taxonomy

This study presents an evidence-informed systematic review and research-readiness taxonomy for culturally aware Malay-English hate speech detection and critically evaluates existing studies based on dataset availability, code-mix authenticity, annotation practice, cultural sensitivity, model architecture, evaluation st...

F. Azmi, Normaisharah Mamat, Rawad Abdulghafor et al. · 0 citations
Open access Aug 2026

Indonesian Hate Speech Detection Across Diverse Domains Using Parameter-Efficient Fine-Tuning with IndoBERT and LoRA

The findings indicate that IndoBERT+LoRA provides a promising and resource-efficient approach for multi-domain Indonesian hate and abusive speech classification, while stricter leave-one-domain-out evaluation remains an important direction for future work.

Fergie Joanda Kaunang, Bhustomy Hakim, A. P. Thenata · 0 citations
Open access 2026

From Binary to Multi-Class: LLM-Judged Synthetic Annotation Applied to Hate Speech Detection

The proposed strategy offers a versatile solution for nuanced classification tasks beyond hate speech, providing a valuable technique for detailed categorisation in various domains.

Antonio Moreno-Cediel, Antonio Garcia-Cabot, Eva García-López · 0 citations
Open access Aug 2026

A Context-Aware and Target-Adaptive Multilingual Framework for Hate Speech Detection in Code-Switched Social Media Text

A Context-Aware and Target-Adaptive Multilingual Hate Speech Detection model that combines multilingual transformer-based embeddings with a context-aware attention mechanism to capture semantic dependencies in text and reduces false positives is introduced.

K. Shruthi, K. Shivanna · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.