Enhancing dialectal Arabic aspect based sentiment analysis through a novel Algerian dialect telecommunication dataset using a unified end-to-end approach
Aug 2026· Journal of King Saud University: Computer and Information Sciences· Vol 38· 0 citations· 59 references
TL;DR
A unified End-to-End (E2E) framework for Arabic ABSA based on the newly introduced dataset and the Arabic ABSA hotels dataset is proposed, where aspect term extraction and sentiment classification are integrated into a single sequence labeling task.
Abstract
This study introduces a novel, publicly available balanced Algerian Arabic Aspect Based Sentiment Analysis (ABSA) dataset consisting of 11,338 comments written in the Algerian dialect. The dataset follows the semantic evaluation 2016 annotation guidelines and includes both explicit and implicit aspects within the telecommunications domain. Furthermore, the study proposes a unified End-to-End (E2E) framework for Arabic ABSA based on the newly introduced dataset and the Arabic ABSA hotels dataset, where aspect term extraction and sentiment classification are integrated into a single sequence labeling task. Transfer learning was leveraged by fine-tuning the Arabic Bidirectional Encoder Representations from Transformers (AraBERT) model, which was further enhanced with Bidirectional Gated Recurrent Units (BiGRU) and a Softmax output layer, forming the fine-tuned AraBERT-BiGRU-Softmax model. Using the unified E2E ABSA approach, the proposed model achieved an overall accuracy of 88.07% on our dataset, along with a macro precision of 72.96%, a macro recall of 60.79%, and a macro F1-score of 65.48% across all labels. When excluding the O label, the model obtained a micro-precision of 54.94%, a micro-recall of 53.85%, and a micro-F1 score of 54.39%. Evaluated on the Arabic ABSA hotel reviews dataset, the model obtained an overall accuracy of 92.91%, with a macro precision of 57.79%, a macro recall of 47.04%, and a macro F1-score of 50.69% across all labels. In addition, when excluding the O label, it reached a micro-precision of 64.74%, a micro-recall of 55.74%, and a micro-F1 score of 59.90%.
This survey introduced a comprehensive systematic review of ASA research from 2018 to 2025, analyzing over 70 peer-reviewed studies and proposing future directions including cross-lingual transfer learning, multimodal sentiment analysis, domain-specific ASA applications.
Ola Adnan Altiti· International journal of com...· 0 citations
This work shows that specialization can outweigh scale in the case of small, low-resourced dialects and highlights the significance of investments into gathering resources for them.
A Hybrid VADER–IndoBERT framework designed to improve sentiment classification robustness on complex Indonesian texts is introduced, demonstrating the superiority of Transformer-based architectures in capturing long-range dependencies and handling ambiguous sentiment cues.
Margareta Valencia Suci Handayani, R. S. Basuki, Muljono et al.· Jurnal RESTI (Rekayasa Siste...· 0 citations
This study successfully proposes a Long Short-Term Memory (LSTM)-based model for automatic classification of Indonesian regional song lyrics by language, demonstrating that LSTM effectively captures sequential linguistic patterns and contextual relationships within regional languages.
Muhammad Rizky, Anandita Priatama, Aviv Yuniar Rahman et al.· Buana Information Technology...· 0 citations
DAMSE is validated on two linguistically distinct datasets and establishes the first zero-shot baselines for vishing detection, introducing dialect-adaptive ensemble weighting that provides a consistent gain over uniform fusion, and releasing the first comprehensive multi-dialect Arabic vishing dataset with structured...
Mohammed Tawfik, A. M. Al-madani, Eman Abdulrahman Alkhamali et al.· PeerJ Computer Science· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.