Aug 2026· JOURNAL OF APPLIED INFORMATICS AND COMPUTING· Vol 10, pp. 3318-3324· 0 citations· 25 references
TL;DR
This practical implementation of the Multinomial Naïve Bayes algorithm combined with Term Frequency-Inverse Document Frequency feature extraction to classify news articles as either factual or hoax provides an accessible and efficient initial screening tool for the general public and journalists to assist in verifying news authenticity.
Abstract
The rapid advancement of information technology in Indonesia has led to a massive spread of digital disinformation, commonly known as an infodemic. The inability to filter inaccurate information manually necessitates a reliable, automated hoax detection system. This study aims to implement and evaluate the Multinomial Naïve Bayes algorithm combined with Term Frequency-Inverse Document Frequency (TF-IDF) feature extraction to classify news articles as either factual or hoax. The research utilizes a dataset of 2,910 Indonesian news articles published in 2025, collected from verified national news portals and fact-checking websites. The text data underwent comprehensive preprocessing—including case folding, cleansing, stopword removal, and stemming—before being evaluated using 5-Fold Cross-Validation and an 80:20 data split. Experimental results demonstrate that the Naïve Bayes model achieves highly stable and competitive performance, recording an accuracy of 93.81%, a precision of 93.84%, a recall of 93.81%, an F1-Score of 93.82%, and a 5-Fold Cross-Validation F1-Score of 93.39%. Notably, the algorithm exhibited a significantly low False Negative rate, missing only 15 hoax documents out of 582 test samples. Furthermore, the trained model was successfully integrated into a real-time, web-based user interface using Streamlit. This practical implementation provides an accessible and efficient initial screening tool for the general public and journalists to assist in verifying news authenticity, thereby supporting efforts to mitigate the impact of digital hoaxes.
Purpose – This study addresses the escalation of disinformation during the 2024 Indonesian General Election by developing an automated hoax detection system. The primary focus is to evaluate the integration of data balancing methods to minimize detection failures in hoax narratives, which often appear less frequently t...
Juliawati Haribae, I. R. H. T. Tangkawarow, G. C. Rorimpandey· Journal of Vocational, Infor...· 0 citations
The machine learning and NLP methods presented in this paper prove that they have the capability to identify misleading news, and this work provides a starting point for machine learning and NLP methods in fictitious news detection.
R. B, S. C, Udayakumar C· International Journal of Sci...· 0 citations
The spread of hoax news in Indonesian online media has become an increasing concern, highlighting the need for accurate automatic detection systems. This study proposes an integrated hoax detection approach that combines IndoBERT, which captures semantic context in Indonesian text, with a Knowledge Graph that verifies...
Laily Maulidya, Sigit Wasista, M. Ruswiansari· Indonesian Journal of Comput...· 0 citations
Across Indonesian online platforms, fabricated news spreads faster than fact-checkers can confirm. Because much of the literature relies on resource-intensive deep models, one applied question stays unsettled: which lighter, more transparent classifier best detects Indonesian hoaxes? We assessed four algorithms, Random...
Dedi Irawan, Sudarmaji· Journal of Information Syste...· 0 citations
Automated fake-news detection is increasingly required because the volume and speed of online content exceed the capacity of manual verification. This study presents a controlled comparison between three conventional machine-learning classifiers---Naive Bayes, Logistic Regression, and Random Forest---and a fine-tuned D...
Abhishek Babu Kolati, Persis Voola· Journal of Computing and Dat...· 0 citations