Aug 2026· Jurnal Minfo Polgan· 0 citations· 23 references
TL;DR
The findings indicate that IndoBERT+LoRA provides a promising and resource-efficient approach for multi-domain Indonesian hate and abusive speech classification, while stricter leave-one-domain-out evaluation remains an important direction for future work.
Abstract
The rapid proliferation of digital connectivity in Indonesia has catalyzed an unprecedented surge in harmful online content, necessitating robust automated systems for hate speech detection that can generalize across diverse digital platforms. Traditional models often struggle with domain shift and the linguistic complexities of Indonesian social media discourse, including informal slang and code-mixing. This research proposes a multi-domain detection framework leveraging the IndoBERT-base-p1 architecture integrated with Low-Rank Adaptation (LoRA), a parameter-efficient fine-tuning (PEFT) strategy. The study utilizes a multi-source corpus from Instagram, Twitter, and news portals, employing back-translation to augment scarce Instagram data and stratified downsampling to ensure domain equilibrium. By training only 1–2% of the total 110 million parameters, specifically targeting the query and value attention modules, the model achieves significant computational savings with a training loss of 0.345. Experimental results demonstrate high robustness, with the framework attaining F1-scores of 0.83 for both Instagram and Twitter, and 0.81 for news portals, while maintaining accuracies between 0.81 and 0.85. Qualitative validation through word cloud analysis further confirms the model's ability to distinguish between aggressive sociopolitical triggers and neutral functional discourse. This study contributes a scalable and resource-efficient solution for real-time content moderation, proving effective across both formal journalistic Indonesian and informal digital dialects. The findings indicate that IndoBERT+LoRA provides a promising and resource-efficient approach for multi-domain Indonesian hate and abusive speech classification, while stricter leave-one-domain-out evaluation remains an important direction for future work.
Due to the widespread accessibility of the internet and social media, toxic and hateful con-tent has grown exponentially, causing significant distress and negative societal impacts. Ro-man Urdu, a low-resource language used in Pakistan and among Urdu-speaking communities worldwide, presents additional challenges becaus...
It is challenging to detect hate speech in Low Resource Languages (LRLs) because of the absence of annotated data, the informality of its language structure, and the lack of standardized grammar. A good example of such a challenge is Roman Urdu which is broadly used by South Asians on social media and has a high variat...
Toneema Zubair, Muhammad Asif, F. Kamiran et al.· 0 citations
A Context-Aware and Target-Adaptive Multilingual Hate Speech Detection model that combines multilingual transformer-based embeddings with a context-aware attention mechanism to capture semantic dependencies in text and reduces false positives is introduced.
K. Shruthi, K. Shivanna· Engineering, Technology &...· 0 citations
This work evaluates the efficiency and competitiveness of seven smaller, more accessible LLMs through zero-shot classification with structured instructions, leveraging the expert-annotated test dataset and annotation scheme of the kNOwHATE project to demonstrate a viable, competitive, and low-cost approach that does no...
Mauro Cardoso, Eugénio Ribeiro, F. Batista et al.· 0 citations
Automated hate speech detection in Indonesian social media remains a persistent challenge due to dataset fragmentation, heterogeneous annotation schemes, and the lack of reproducible cross-model benchmarks with formal statistical validation. This study presents a cross-validated benchmark that systematically evaluates...
Dodo Zaenal Abidin, Agus Siswanto, Chindra Saputra et al.· Jurnal Teknik Informatika (J...· 0 citations
Online company review platforms have gained significant importance as they provide transparent insights into corporate culture and employee satisfaction. However, analyzing this feedback at scale remains challenging due to the linguistic complexity of workplace-specific discourse and the lack of high-quality, multi-dim...
Khanh-Long Ho-Vuong, Nhat-Huy Dang, Do Bao et al.· International Conference on...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.