A Comparative Analysis of Large Language Models for the Detection and Classification of Hate Speech in a Low-Resource Language
TL;DR
The presented results highlight the complexity and diversity of hate speech in Serbian online communication, demonstrating a high detection accuracy of 86.7% achieved with the LLaMA 3 model, followed by Qwen3 (82%) for the two-stage sentence-level pipeline and Qwen3 (82%) for the two-stage sentence-level pipeline.
Abstract
Hate speech represents one of the most significant challenges of modern digital society, particularly due to the pervasive use of social networks and online media. Developing reliable automated detection systems requires high-quality, meticulously curated, and annotated datasets - a task that is especially challenging for low-resource languages, such as Serbian. This paper describes the process of collecting, processing, and labeling textual data aimed at creating a dataset for hate speech detection in the Serbian language, as well as a comparative analysis of Large Language Models (LLMs) in the detection and classification of hate speech. The data were gathered from diverse sources, including social networks and online media platforms, utilizing both automated and manual techniques, as well as web crawling and scraping methods. The resulting dataset comprises 1351 short texts, containing 8029 sentences, annotated into three distinct classes: non-hate speech, offensive speech, and hate speech. Furthermore, hate speech instances were additionally categorized according to relevant types of discrimination in accordance with European Union legal acts. In addition to describing the annotation process, this paper provides a detailed analysis of the dataset, including class distribution, text length, and the most frequent keywords. Subsequently, the research selected five LLMs, which were inferred using prompt engineering with a different design for each approach. The models were evaluated using a literature-informed zero-shot prompting strategy based on detailed category definitions, decision criteria, and structured output constraints, implemented through single-stage, two-stage, and paragraph-level prompting settings. The selected models represent recent and widely used open-source general-purpose LLMs available through the Ollama platform, which was chosen to enable local, reproducible, and privacy-preserving evaluation on consumer hardware. Additional comparative experiments were also conducted using few-shot prompting, a general-purpose LLM, and the supervised BERTić model, and the publicly available bcms-bertic-frenk-hate classifier as external baselines. These LLMs were then employed for the detection and classification of hate speech, utilizing an ensemble strategy. The presented results highlight the complexity and diversity of hate speech in Serbian online communication, demonstrating a high detection accuracy of 86.7% achieved with the LLaMA 3 model, followed by Qwen3 (82%) for the two-stage sentence-level pipeline. For paragraph-level analysis, Qwen achieved the highest detection accuracy of 80%, slightly better than LLaMA 3 (77.3%). Ensemble of all five LLMs achieved comparable results to the best selected models, achieving 81.4% in two-stage sentence-level and 79.1% in paragraph-level analysis.