A Comparative Analysis of Large Language Models for the Detection and Classification of Hate Speech in a Low-Resource Language
The presented results highlight the complexity and diversity of hate speech in Serbian online communication, demonstrating a high detection accuracy of 86.7% achieved with the LLaMA 3 model, followed by Qwen3 (82%) for the two-stage sentence-level pipeline and Qwen3 (82%) for the two-stage sentence-level pipeline.