Low-Resource Hate Speech Detection in English-Swahili Code-Switched Text Using Fine-Tuning of Pre-trained Language Models
The use of social media in East Africa has grown rapidly, and with it, the spread of hate speech has become a serious concern. This problem is even more complex in online spaces where people often switch between English and Swahili within the same sentence or conversation. Such code-switching makes it difficult for existing systems to accurately detect harmful content, especially because there is limited labeled data and much of the language used is informal and context-dependent. This study explores a low-resource approach to detecting hate speech in English and Swahili code-switched text by fine-tuning pre-trained language models. In this work, transformer-based models such as BERT and AfriBERTa are adapted to better understand mixed-language communication. The models are trained on a carefully prepared dataset made up of real social media posts that reflect how people actually write and speak online. These posts are manually labeled to capture both direct and subtle forms of hate speech, including expressions that are influenced by local culture and everyday slang. The findings show that fine-tuned models perform better than traditional machine learning approaches, especially in terms of accuracy and overall detection quality. They are also more effective at handling informal language, abbreviations, and mixed grammar structures. Beyond performance, the study also looks at fairness and bias, emphasizing the need for systems that are sensitive to cultural and linguistic diversity. Overall, this work shows that fine-tuning modern language models can offer a practical and scalable solution for hate speech detection in multilingual environments.