Skip to content
Open access

Comparison of Shallow and Deep Learning for Indonesian Clickbait Headline Classification

Jul 2026 · Journal of Information System Exploration and Research · Vol 4, pp. 281-290 · 0 citations · 28 references

TL;DR

This study contributes to Indonesian clickbait detection research by demonstrating that ensemble aggregation of diverse transformer architectures yields more reliable performance than reliance on any single model.

Abstract

Clickbait is an increasingly prevalent phenomenon in Indonesian online news media, where headlines are crafted to attract clicks without accurately reflecting article content. This study proposes and compares eight classification models: three shallow learning algorithms — Naive Bayes, Support Vector Machine (SVM), and Logistic Regression — and five transformer-based deep learning models: IndoBERT-p1, IndoBERT-p2, XLM-RoBERTa, mBERT, and DistilBERT. The dataset used is CLICK-ID, consisting of 15,000 labeled headlines from 12 Indonesian news portals, expanded to 25,138 samples via semi-supervised pseudo labeling with a confidence threshold of 0.85. All deep learning models were trained with Focal Loss (α=0.25, γ=2.0) to address class imbalance and Automatic Mixed Precision (AMP) for GPU efficiency. Results show that IndoBERT-p1, IndoBERT-p2, XLM-RoBERTa, mBERT, and DistilBERT achieve comparable performance, with macro F1-scores ranging from 86.57% to 88.54%. Among shallow learning models, SVM performs best with 83.51% F1-score. An average ensemble of all five transformer models achieves the best overall performance at 90.14% accuracy and 89.00% F1-score, outperforming every individual model. This study contributes to Indonesian clickbait detection research by demonstrating that ensemble aggregation of diverse transformer architectures yields more reliable performance than reliance on any single model.

Read PDF

Similar papers

Conference Aug 2026

Fake News Detection Using Deep Learning with Continual Learning

Misinformation propagation across online platforms continues to pose serious risks to informed public discourse and media credibility. To address this, we design and evaluate a fully integrated fake news detection pipeline built upon the FakeNewsNet benchmark, drawing from both PolitiFact and Buz-zFeed corpora. This wo...

G. Sai, R. B. Kumar, Yalavarthi Sai Eswari · 0 citations
Open access Aug 2026

The First Comprehensive Study of Stance Detection Modeling for the Sorani Kurdish Language

Stance detection has become a fundamental task in natural language processing (NLP), yet it remains under-explored for low-resource languages such as Sorani Kurdish. Building on the previously released Bochun dataset, the present work focuses exclusively on the comprehensive evaluation of stance detection models and pr...

P. S. Rostam, Rebwar M. Nabi · 0 citations
Open access Aug 2026

Comparative Analysis of Four Machine Learning Classifiers for Indonesian Hoax News Detection

Across Indonesian online platforms, fabricated news spreads faster than fact-checkers can confirm. Because much of the literature relies on resource-intensive deep models, one applied question stays unsettled: which lighter, more transparent classifier best detects Indonesian hoaxes? We assessed four algorithms, Random...

Dedi Irawan, Sudarmaji · 0 citations
Open access Aug 2026

AugLog-LightGBM: A Log-Based Feature AugmentationFramework for Class Imbalance in Credit RiskClassification

Non-performing loan (NPL) detection is inherently a class-imbalance problem because defaulting borrowers represent a persistent minority. Standard gradient boosting often favors the majority class. This paper proposes AugLog-LightGBM, an extension of LightGBM that improves initialization through Log-Based Feature Augme...

H. Azizah, E. Sumarminingsih, A. Fernandes · 0 citations
Review Open access Aug 2026

Sentiment Analysis of Amazon Mobile Reviews Using Deep Learning Techniques for Brand Performance Evaluation

Nowadays, Natural Language Processing, or NLP, is a key component of many programs that analyze and comprehend human language. The sentiment analysis of mobile product reviews collected from the Kaggle repository—more especially, the 20,710-review Amazon Mobile evaluations dataset—is the main emphasis of this research....

Dhananchezhiyan R, M. Rameshkumar · 0 citations
Open access Sep 2026

Deep Learning Methods for Multimodal Fake News Classification Combining Textual and Visual Information

The authors suggest a computationally efficient multimodal deep learning framework using Bidirectional Encoder Representations of Transformers (BERT) to extract textual features and convolutional neural networks to learn visual representations that offers a computational scaling alternative to attention-based models, w...

P. Jadhav, R. K. Shukla · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.