Skip to content
Open access

From Binary to Multi-Class: LLM-Judged Synthetic Annotation Applied to Hate Speech Detection

2026 · Computers, Materials & Continua · Vol 89, pp. 1-10 · 0 citations · 83 references

TL;DR

The proposed strategy offers a versatile solution for nuanced classification tasks beyond hate speech, providing a valuable technique for detailed categorisation in various domains.

Abstract

: The increasing prevalence of hate speech on social media platforms has spurred research aimed at mitigating this societal harm. However, the development of effective machine learning solutions is hindered by a lack of labelled hate speech data in languages beyond English, particularly when attempting granular, multi-class classification. This research aims to address this data scarcity by introducing a novel methodology leveraging the ‘Large Language Model as a judge’ paradigm to transform existing binary-labelled hate speech data into multi-class datasets. Our approach aims to generate balanced datasets and enables classification across seven identity groups: race, religion, origin, gender, sexuality, age, and disability. The methodology has been applied to the Spanish Hate Speech Superset, and it has been validated using the Measuring Hate Speech dataset, demonstrating significant efficacy and broad applicability. Specifically, our approach obtains a higher match rate with human labels and a lower number of mismatches when compared with prompt-only strategies. In addition, Cohen’s Kappa scores demonstrate that our approach exhibits a moderate strength of agreement with human annotators, outperforming prompt-only strategies’ scores by 5%. As a result of the application of the proposed strategy to the Spanish Hate Speech Superset dataset, a multi-class version is obtained, comprising 6325 hateful samples classified among the seven identity groups. The proposed strategy offers a versatile solution for nuanced classification tasks beyond hate speech, providing a valuable technique for detailed categorisation in various domains.

Read PDF

Similar papers

Cost-Effective Hate Speech Detection in Portuguese Using Lightweight LLMs

This work evaluates the efficiency and competitiveness of seven smaller, more accessible LLMs through zero-shot classification with structured instructions, leveraging the expert-annotated test dataset and annotation scheme of the kNOwHATE project to demonstrate a viable, competitive, and low-cost approach that does no...

Mauro Cardoso, Eugénio Ribeiro, F. Batista et al. · 0 citations
Open access Sep 2026

Fine-tuning large language models for binary and multiclass target-specific hate speech detection

Social media platforms are widely used by individuals from diverse demographic backgrounds to share their opinions and daily experiences. The anonymity provided by these platforms has increased the prevalence of hate speech, causing significant psychological harm to targeted groups. Existing hate speech detection appro...

Sanaa Kaddoura, Sumaia A. Al-Kohlani · 0 citations
Open access Aug 2026

Indonesian Hate Speech Detection Across Diverse Domains Using Parameter-Efficient Fine-Tuning with IndoBERT and LoRA

The findings indicate that IndoBERT+LoRA provides a promising and resource-efficient approach for multi-domain Indonesian hate and abusive speech classification, while stricter leave-one-domain-out evaluation remains an important direction for future work.

Fergie Joanda Kaunang, Bhustomy Hakim, A. P. Thenata · 0 citations
#computer vision Preprint Sep 2026

MexHat: A Dataset for Hate Speech Detection in Mexican Spanish Videos

Ensuring online safety through content monitoring had raised Hate Speech Detection as a crucial task to be addressed. By essence the task demands the capture of contextual cues, which are essential for a precise understanding of the content's intent. Although automated detection approaches for the task have advanced si...

Itzel Tlelo-Coyotecatl, Hugo Jair Escalante · 0 citations
Open access Aug 2026

A Context-Aware and Target-Adaptive Multilingual Framework for Hate Speech Detection in Code-Switched Social Media Text

A Context-Aware and Target-Adaptive Multilingual Hate Speech Detection model that combines multilingual transformer-based embeddings with a context-aware attention mechanism to capture semantic dependencies in text and reduces false positives is introduced.

K. Shruthi, K. Shivanna · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.