Skip to content
Open access

Harnessing Ensemble and Transformers for Sentiment Analysis and Emotion Detection in Hausa Text

2026 · International journal of research and innovation in social science · Vol 10, pp. 15860-15876 · 0 citations

TL;DR

This research establishes a rigorous benchmark for Hausa NLP, highlighting the indispensable role of subword tokenization, contextual embeddings, and threshold calibration in handling the morphological richness and multi-label complexities of low-resource languages.

Abstract

Understanding emotional tone and sentiment in text has driven significant advancements in Natural Language Processing (NLP), particularly in sentiment analysis and emotion detection. This study addresses the challenge of developing effective NLP tools for low-resource languages, focusing on the Hausa language. By leveraging ensemble methods and pre-trained transformer models like BERT and XLM-R, along with traditional classifiers such as Logistic Regression, SVM, Naive Bayes, Random Forest, and XGBoost, we aim to improve sentiment analysis and emotion detection for Hausa text. Utilizing a balanced sentiment dataset (9,958 samples) and a complex multi-label emotion dataset (19,757 samples across 11 categories), we benchmark individual classifiers, voting ensembles, and deep contextual models. For sentiment analysis, a Hard Voting Ensemble of TF-IDF-vectorized base learners achieved a highly competitive F1-score of 0.8748. However, Transformer models significantly outperformed traditional baselines, with Multilingual BERT (mBERT) achieving a peak F1-score of 0.8983. In the multi-label emotion detection task, individual traditional models struggled with label sparsity, yielding low Subset Accuracy scores (2.88% to 8.30%) and moderate Micro-F1 scores. Standard Hard Voting ensembles further underperformed due to discrete prediction conflicts. To resolve this, a Probability-based Majority Voting mechanism with calibrated thresholding (0.3) was introduced, boosting the Micro-F1 to 0.3825 and reducing the Hamming Loss to 0.1967. Ultimately, XLM-RoBERTa emerged as the superior architecture, achieving a Subset Accuracy of 0.1545, a Micro-F1 of 0.4275, and the lowest Hamming Loss of 0.1804. This research establishes a rigorous benchmark for Hausa NLP, highlighting the indispensable role of subword tokenization, contextual embeddings, and threshold calibration in handling the morphological richness and multi-label complexities of low-resource languages

Read PDF

Similar papers

Review Open access 2020

Improving Sentiment Analysis Using Transformer-Based NLP Models

Experimental results on benchmark datasets show that transformer models outperform traditional methods in accuracy, precision, recall, and F1-score, highlighting that transformer-based approaches provide more efficient and scalable solutions for real-world sentiment analysis applications.

Ibrahim Lawal · 0 citations
Open access 2026

A Hybrid Framework for Large-Scale Tweet Sentiment Analysis Using Classical Machine Learning, Transformer Models, and Uncertainty Estimation

The work provides a reproducible, explainable, operationally applicable model of sentiment analysis in operationally sensitive, high-stakes Twitter sentiment analysis, and validate the hypothesis that hybrid stacking is an effective method for leveraging the complementary nature of lexical and contextual representation...

D. Abate, Nilay Mistry · 0 citations
Review Aug 2026

Explainable Artificial Intelligence (AI) In Sentiment Analysis

Experimental results demonstrate that XAI techniques significantly enhance the interpretability of sentiment prediction without substantially compromising classification performance, and explainable sentiment analysis supports fairness assessment, bias detection, regulatory compliance, and informed decision-making in c...

Aishwarya P. A., N. K · 0 citations
Review Open access Aug 2026

OPTIMIZING MACHINE LEARNING CLASSIFIERS FOR HIGH-ACCURACY SENTIMENT DETECTION IN THE TURKISH LANGUAGE

This study investigates the optimization of classical machine learning classifiers and ensemble learning strategies for binary Turkish sentiment analysis under a unified experimental framework and demonstrates that optimized classical models remain highly effective in Turkish SA, and their accuracy can be further impro...

Ahmad Bwidani, Ali A. H. Karah Bash · 0 citations
Review Open access Jul 2026

Enhanced Sentiment Analysis Using RoBERTa and BiLSTM: A Context-Aware Hybrid Deep Learning Approach

This paper presents a context-aware hybrid deep learning approach by integrating the Robustly Optimized BERT Pretraining Approach (RoBERTa) with Bidirectional Long Short-Term Memory (BiLSTM) networks to generate rich contextual word embeddings.

V. Gayatri, Rajani Rajalingam · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.