Skip to content
Open access

Comparative Evaluation of IndoBERT-Based Architectures for Imbalanced Indonesian News Title Classification

Jul 2026 · SinkrOn · Vol 10, pp. 1896-1907 · 1 citation · 19 references

TL;DR

Empirical evidence is shown indicating that composing several imbalance-oriented techniques on a pretrained transformer can yield adverse interactions rather than cumulative gains.

Abstract

Stacking multiple imbalance-mitigation techniques on top of a pretrained transformer is widely assumed to compound their individual benefits, yet rigorous component-wise evidence for this assumption remains scarce in the Indonesian text classification literature. Four classification architectures are compared in this work on a publicly available Indonesian news title corpus. The working set contains 27,266 short headlines, drawn as a 30% stratified subsample from a cleaned corpus of 90,891 headlines, spread over nine target categories with a class ratio of 13.29. Three reference architectures are constructed: an LSTM trained from scratch with Random Oversampling, a bidirectional LSTM augmented with additive attention, and a fine-tuned IndoBERT on the oversampled training partition. A fourth architecture extends IndoBERT through three additions, namely learned attention pooling over contextual token embeddings, focal modulation applied on top of the cross-entropy term, and minority-class paraphrasing via Indonesian–English–Indonesian back-translation. Every configuration is evaluated through stratified 5-fold cross-validation, paired t-tests with Bonferroni correction across three comparisons, and McNemar tests on the held-out partition. The fine-tuned IndoBERT with Random Oversampling alone reaches the highest macro F1 of 0.837. By contrast, the combined configuration drops to 0.799, and statistical verification confirms that the gap is systematic rather than attributable to fold-level variation. A component-wise ablation isolates focal modulation as the principal driver of the decline, because it disturbs an already-balanced training distribution. The principal outcome of this study is empirical evidence indicating that composing several imbalance-oriented techniques on a pretrained transformer can yield adverse interactions rather than cumulative gains.

Read PDF

Similar papers

Aug 2026

LiteLLM: a lightweight transformer architecture for efficient short-text classification

LiteLLM is introduced, a lightweight transformer architecture explicitly optimized for short-text scenarios that delivers competitive performance, fast convergence, and competitive cross-domain performance across heterogeneous short-text settings.

Hussein Ala’a Alkaabi, Fuqdan A. Al-Ibraheemi, Ali kadhim Jasim · 0 citations
Open access Jul 2026

Comparison of Shallow and Deep Learning for Indonesian Clickbait Headline Classification

This study contributes to Indonesian clickbait detection research by demonstrating that ensemble aggregation of diverse transformer architectures yields more reliable performance than reliance on any single model.

Muhammad Noer Attalah Dzahkwan, Majid Rahardi · 0 citations
Preprint Aug 2026

AraSSM: A bidirectional state-space encoder for Arabic masked language modeling

A bidirectional Mamba encoder pretrained via masked language modeling on a corpus combining Arabic Wikipedia and CulturaX text is introduced, trained end-to-end on four consumer-grade NVIDIA RTX 2080Ti GPUs (11GB) over approximately ten days.

A. A. Aliane, H. Aliane, N. Semmar · 0 citations
Review Open access Jul 2026

A Hybrid VADER–IndoBERT Framework for Robust Sentiment Analysis of Long and Ambiguous Indonesian Texts

A Hybrid VADER–IndoBERT framework designed to improve sentiment classification robustness on complex Indonesian texts is introduced, demonstrating the superiority of Transformer-based architectures in capturing long-range dependencies and handling ambiguous sentiment cues.

Margareta Valencia Suci Handayani, R. S. Basuki, Muljono et al. · 0 citations
Open access Sep 2026

A synthetic minority over-sampling–enhanced hybrid Indonesian bidirectional encoder representations from transformers–convolutional neural network model with cross-entropy optimization for imbalanced fake news detection in Indonesian online media

This study introduces a hybrid Indonesian Bidirectional Encoder Representations from Transformers-Convolutional Neural Network (IndoBERT-CNN) model enhanced through cross-entropy optimization, Word2Vec-based embedding refinement, and systematic imbalanced data handling using several Synthetic Minority Over-sampling var...

Yuliant Sibaroni, S. Prasetiyowati, Diyas Puspandari · 0 citations
Open access Jul 2026

LSTM-Based Classification of Indonesian Regional Song Lyrics by Language

This study successfully proposes a Long Short-Term Memory (LSTM)-based model for automatic classification of Indonesian regional song lyrics by language, demonstrating that LSTM effectively captures sequential linguistic patterns and contextual relationships within regional languages.

Muhammad Rizky, Anandita Priatama, Aviv Yuniar Rahman et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.