Skip to content
Review Open access

Culturally Aware Malay–English Code-Mixed Hate Speech Detection: A Systematic Review and Research Taxonomy

2026 · IEEE Access · Vol 14, pp. 128979-129003 · 0 citations · 79 references

TL;DR

This study presents an evidence-informed systematic review and research-readiness taxonomy for culturally aware Malay-English hate speech detection and critically evaluates existing studies based on dataset availability, code-mix authenticity, annotation practice, cultural sensitivity, model architecture, evaluation strategy, explainability, robustness and deployment readiness.

Abstract

Social media platforms such as X, formerly Twitter, have become major channels for communication, information sharing and public discussion. However, the rapid growth of user-generated content has also increased the spread of hate speech and offensive language. Automated hate speech detection remains challenging in multilingual and code-mixed environments, where users frequently combine languages, informal spelling, slang, abbreviations and culturally specific expressions. In Malaysia, online discourse often involves Malay-English code-mixing, commonly referred to as Manglish, which creates additional challenges for natural language processing systems. This study presents an evidence-informed systematic review and research-readiness taxonomy for culturally aware Malay-English hate speech detection. Unlike conventional reviews that mainly summarize model architecture and performance scores, this review critically evaluates existing studies based on dataset availability, code-mix authenticity, annotation practice, cultural sensitivity, model architecture, evaluation strategy, explainability, robustness and deployment readiness. To strengthen this review, this article incorporates a completed empirical case study on Manglish hate-speech detection, using posts collected from X (formerly Twitter). The case study used keyword-based data collection, Malaya NLP-based language filtering, bilingual manual annotation, manual class balancing and transformer-based evaluation using mBERT, XLNet and XLM-RoBERTa. The case evidence is used only as an empirical lens to illustrate practical challenges in dataset curation, class imbalance, lexical overlap and model robustness. It is not positioned as a new benchmark dataset or an independent experimental contribution. The review finds that existing Malay and Malay-English resources remain fragmented. Some datasets are monolingual Malay hate speech datasets, some are bilingual but language-separated, while others are code-mixed but developed for sentiment analysis rather than hate speech detection. Transformer-based models such as BERT, mBERT and XLM-RoBERTa show strong potential, but their results are difficult to compare due to inconsistent datasets, label definitions, class distributions, evaluation metrics and limited robustness testing. The main novelty of this review is the proposed taxon- omy that evaluates the field through six readiness dimensions, which are data, linguistic, cultural, modelling, evaluation and deployment readiness. The findings from this study provide a structured foundation for developing culturally aware, robust, explainable and deployable hate speech detection systems for Malay-English code-mixed social media.

Read PDF

Similar papers

Review Open access Jul 2026

Detecting Hate Speech in Hindi Digital Discourse Using Transformer–Long Short-Term Memory Models

Evaluating advanced language models for analysing hate speech in Hindi social media content illustrates that language-specific computational tools can be used for both platform governance and communication research, provided that the cultural context is considered.

Rachna Narula, Vedika Gupta, Jawad Khan et al. · 0 citations
Open access Jul 2026

Low-Resource Hate Speech Detection in English-Swahili Code-Switched Text Using Fine-Tuning of Pre-trained Language Models

This study explores a low-resource approach to detecting hate speech in English and Swahili code-switched text by fine-tuning pre-trained language models, and shows that fine-tuning modern language models can offer a practical and scalable solution for hate speech detection in multilingual environments.

Kipkebut Andrew, Jepkemei Betty · 0 citations
Open access Aug 2026

A Context-Aware and Target-Adaptive Multilingual Framework for Hate Speech Detection in Code-Switched Social Media Text

A Context-Aware and Target-Adaptive Multilingual Hate Speech Detection model that combines multilingual transformer-based embeddings with a context-aware attention mechanism to capture semantic dependencies in text and reduces false positives is introduced.

K. Shruthi, K. Shivanna · 0 citations
Open access 2026

A Comparative Analysis of Large Language Models for the Detection and Classification of Hate Speech in a Low-Resource Language

The presented results highlight the complexity and diversity of hate speech in Serbian online communication, demonstrating a high detection accuracy of 86.7% achieved with the LLaMA 3 model, followed by Qwen3 (82%) for the two-stage sentence-level pipeline and Qwen3 (82%) for the two-stage sentence-level pipeline.

Matija Dodović, Janko Tufegdžić, Dražen Drašković · 0 citations
Open access Aug 2026

Indonesian Hate Speech Detection Across Diverse Domains Using Parameter-Efficient Fine-Tuning with IndoBERT and LoRA

The findings indicate that IndoBERT+LoRA provides a promising and resource-efficient approach for multi-domain Indonesian hate and abusive speech classification, while stricter leave-one-domain-out evaluation remains an important direction for future work.

Fergie Joanda Kaunang, Bhustomy Hakim, A. P. Thenata · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.