Skip to content
Open access

A synthetic minority over-sampling–enhanced hybrid Indonesian bidirectional encoder representations from transformers–convolutional neural network model with cross-entropy optimization for imbalanced fake news detection in Indonesian online media

Sep 2026 · PeerJ Computer Science · Vol 12, pp. e4080 · 0 citations · 54 references

TL;DR

This study introduces a hybrid Indonesian Bidirectional Encoder Representations from Transformers-Convolutional Neural Network (IndoBERT-CNN) model enhanced through cross-entropy optimization, Word2Vec-based embedding refinement, and systematic imbalanced data handling using several Synthetic Minority Over-sampling variants.

Abstract

The dissemination of fake news in Indonesian online media, particularly in political discourse, continues to increase and is often characterized by highly imbalanced class distributions. This study proposes a robust detection framework tailored for political fake news in the Indonesian language, with potential adaptability to other languages through appropriate linguistic and domain adjustments. The framework integrates a series of data mining and deep learning optimization strategies, emphasizing performance stability under severe data imbalance conditions. Specifically, we introduce a hybrid Indonesian Bidirectional Encoder Representations from Transformers-Convolutional Neural Network (IndoBERT-CNN) model enhanced through cross-entropy optimization, Word2Vec-based embedding refinement, and systematic imbalanced data handling using several Synthetic Minority Over-sampling (SMOTE) variants. Model performance was rigorously assessed using 10-fold cross-validation and benchmarked against multiple baseline architectures, including Convolutional Neural Network (CNN), Long Short-Term Memory (LSTM), Indonesian Bidirectional Encoder Representations from Transformers (IndoBERT), and Convolutional Neural Network—Long Short-Term Memory (CNN-LSTM). The experimental results show that the proposed model consistently outperforms all baseline models across various imbalanced scenarios, achieving a best F1-score of 90%, indicating its strong generalization capability and suitability for political fake news detection in Indonesian online media.

Read PDF

Similar papers

Open access Aug 2026

WEIGHTED LOSS STRATEGY FOR BERT-BASED TWITTER SENTIMENT ANALYSIS WITHOUT SYNTHETIC OVERSAMPLING

The widespread adoption of ChatGPT has generated extensive public discourse across social media, necessitating robust sentiment analysis to understand collective opinions. Traditional approaches frequently employ the Synthetic Minority Over-sampling Technique (SMOTE) to address class imbalance; however, its effectivene...

T. Siallagan, R. Winanjaya, Juni Ismail · 0 citations
Open access Aug 2026

Sentiment Analysis of Imbalanced Dataset Through Data Augmentation and Generative Annotation Using DistilBERT and Low‐Rank Fine‐Tuning

Sentiment analysis on social media data often suffers from severe class imbalance, which can negatively affect the performance of machine learning models. In this paper, we propose a framework that leverages large language models and lightweight transformer fine tuning to improve sentiment classification on imbalance...

Hossein Nekkouei Nasrabadi, M. Moattar · 1 citation
Open access Sep 2026

BBCS-Net: A BERT-Based CNN-BiLSTM-SVM Hybrid Architecture for Fine-Grained Cyberbullying Detection

Cyberbullying detection remains challenging due to diverse linguistic patterns used by offenders. The in- ductive biases of deep learning architectures add to this challenge. Single-model approaches often capture only part of abusive language. This results in distinct but partially overlapping error spaces, limiting th...

K. Bhalerao, Vanita Mane · 0 citations
Open access Sep 2026

Web-Integrated Deepfake Detection of Manipulated Political Audio Using CNN and Wav2Vec2 Transformer Models

Deepfake audio poses increasing risks to political forensics by enabling scalable misinformation, speaker impersonation, and identity spoofing. In response to this challenge, this paper presents a web-integrated deepfake detection framework for manipulated political speech together with a politically grounded multiling...

Patricio Mendoza Nuñez, Juan Pablo Astudillo León, Jefferson Alexander Moreno-Guaicha et al. · 0 citations
Review Open access Sep 2026

A COMPARATIVE STUDY OF CNN-BILSTM AND TF-IDF–NAIVE BAYES FOR SENTIMENT CLASSIFICATION ON MOVIE REVIEWS: PERFORMANCE AND EFFICIENCY TRADE-OFFS

Sentiment analysis on movie reviews has become an important task for understanding public opinion, yet many high-performing deep learning models proposed in the literature rely on increasingly complex architectures such as multichannel embeddings, kernel-based projections, or attention-augmented transformers. This adde...

Muhammad Fadhil Dwisaputra · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.