Jul 2026· International Seminar on Intelligent Technology and Its Applications· pp. 541-546· 0 citations· 38 references
Abstract
The lack of high-quality labeled datasets remains a major challenge for sentiment analysis in low-resource languages such as Indonesian, particularly in specialized domains like fiscal policy. This study investigates the effectiveness of Large Language Models (LLMs) as automated annotators within a teacher-student knowledge distillation framework. Using social media data from X related to Indonesia's Coretax system, three training scenarios were evaluated: AI-labeled data, human-labeled data, and a hybrid approach. The results show that GPT-4o achieves substantial agreement with human annotators, with a Cohen's Kappa score of 0.61. Furthermore, the student model IndoBERT trained on the combined dataset outperforms other configurations, achieving a Macro F1-score of 0.64 and a Macro ROC-AUC of 0.84. These findings indicate that while LLMs cannot fully replace human judgment, they significantly enhance scalability and enable near real-time policy evaluation in low-resource settings through effective human-AI collaboration.
The integration of the IndoBERT-BiLSTM architecture with SHAP is demonstrated to deliver accurate and explainable Indonesian sentiment analysis, which effectively bridges the gap between deep learning performance and decision transparency without compromising classification accuracy.
A. Widiyatmoko, A. Nugroho, Muhammad Nurul Firdaus· Journal of Electrical Engine...· 0 citations
Social media sentiment analysis has become one of the most significant instruments for understanding the opinion of the population in the spheres of healthcare, politics, and education. Yet, large language models (LLMs) remain unevenly distributed in their linguistic coverage, failing to adequately serve a large portio...
Muhamet Kastrati, Abdul Manaf, A. Imran et al.· Frontiers in Artificial Inte...· 0 citations
This study aims to analyze public sentiment toward the LPDP alumni controversy on social media using a deep learning approach. The research data consist of YouTube user comments related to the LPDP issue, which were processed through text preprocessing and automatically labeled using IndoBERT into three sentiment class...
Dwi Erzalianti, Joice Junansi Tandirerung, C. Suhaeni et al.· JOURNAL OF APPLIED INFORMATI...· 0 citations
Identifying the target of emotional words or phrases in crisis situations, especially health-related ones, is important for understanding public concerns across cultural and linguistic contexts. We propose CrisisKD, a five-stage teacher--student knowledge distillation framework for aspect-level sentiment and emotion an...
Marko Haralović, Onat Akça, Salih Eren Yücetürk et al.· 0 citations
The results of the experiments reveal that the proposed DistilBERT-based model outperforms the baseline Conditional Random Field and Bidirectional Long Short-Term Memory models with an accuracy 79.05%, precision 80.63%, recall 79.05% and F1-score 79.24%.
Maureen Otieno, L. Wanzare, Calvins Otieno· International Journal of Com...· 0 citations
While sentiment analysis has advanced significantly, fine-grained sentiment classification such as aspect-based sentiment analysis (ABSA), continues to present challenges. These difficulties primarily stem from data scarcity and the inherent complexities of identifying sentiments specific to different aspects within...
Ling-Ling Xu, Hao-Ran Xie, S. Qin et al.· International Conference on...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.