Skip to content
Open access

The Application of Naïve Bayes Algorithm in Detecting Hoaxes on National News Portals in Indonesia

Aug 2026 · JOURNAL OF APPLIED INFORMATICS AND COMPUTING · Vol 10, pp. 3318-3324 · 0 citations · 25 references

TL;DR

This practical implementation of the Multinomial Naïve Bayes algorithm combined with Term Frequency-Inverse Document Frequency feature extraction to classify news articles as either factual or hoax provides an accessible and efficient initial screening tool for the general public and journalists to assist in verifying news authenticity.

Abstract

The rapid advancement of information technology in Indonesia has led to a massive spread of digital disinformation, commonly known as an infodemic. The inability to filter inaccurate information manually necessitates a reliable, automated hoax detection system. This study aims to implement and evaluate the Multinomial Naïve Bayes algorithm combined with Term Frequency-Inverse Document Frequency (TF-IDF) feature extraction to classify news articles as either factual or hoax. The research utilizes a dataset of 2,910 Indonesian news articles published in 2025, collected from verified national news portals and fact-checking websites. The text data underwent comprehensive preprocessing—including case folding, cleansing, stopword removal, and stemming—before being evaluated using 5-Fold Cross-Validation and an 80:20 data split. Experimental results demonstrate that the Naïve Bayes model achieves highly stable and competitive performance, recording an accuracy of 93.81%, a precision of 93.84%, a recall of 93.81%, an F1-Score of 93.82%, and a 5-Fold Cross-Validation F1-Score of 93.39%. Notably, the algorithm exhibited a significantly low False Negative rate, missing only 15 hoax documents out of 582 test samples. Furthermore, the trained model was successfully integrated into a real-time, web-based user interface using Streamlit. This practical implementation provides an accessible and efficient initial screening tool for the general public and journalists to assist in verifying news authenticity, thereby supporting efforts to mitigate the impact of digital hoaxes.

Read PDF

Similar papers

Open access Sep 2026

Implementation of the Multinomial Naïve Bayes Algorithm in a Web-Based System for Detecting Online News Hoaxes During the 2024 Elections

Purpose – This study addresses the escalation of disinformation during the 2024 Indonesian General Election by developing an automated hoax detection system. The primary focus is to evaluate the integration of data balancing methods to minimize detection failures in hoax narratives, which often appear less frequently t...

Juliawati Haribae, I. R. H. T. Tangkawarow, G. C. Rorimpandey · 0 citations
Open access Aug 2026

Detecting the Deception : An Intelligent Machine Fake News Detection

The machine learning and NLP methods presented in this paper prove that they have the capability to identify misleading news, and this work provides a starting point for machine learning and NLP methods in fictitious news detection.

R. B, S. C, Udayakumar C · 0 citations
Open access Aug 2026

Indonesian Hoax News Detection Using IndoBERT and Knowledge Graph

The spread of hoax news in Indonesian online media has become an increasing concern, highlighting the need for accurate automatic detection systems. This study proposes an integrated hoax detection approach that combines IndoBERT, which captures semantic context in Indonesian text, with a Knowledge Graph that verifies...

Laily Maulidya, Sigit Wasista, M. Ruswiansari · 0 citations
Open access Aug 2026

Comparative Analysis of Four Machine Learning Classifiers for Indonesian Hoax News Detection

Across Indonesian online platforms, fabricated news spreads faster than fact-checkers can confirm. Because much of the literature relies on resource-intensive deep models, one applied question stays unsettled: which lighter, more transparent classifier best detects Indonesian hoaxes? We assessed four algorithms, Random...

Dedi Irawan, Sudarmaji · 0 citations
Open access Sep 2026

A Comparative Framework for Fake News Detection Using DistilBERT and Machine Learning Algorithms

Automated fake-news detection is increasingly required because the volume and speed of online content exceed the capacity of manual verification. This study presents a controlled comparison between three conventional machine-learning classifiers---Naive Bayes, Logistic Regression, and Random Forest---and a fine-tuned D...

Abhishek Babu Kolati, Persis Voola · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.