A Comparative Framework for Fake News Detection Using DistilBERT and Machine Learning Algorithms
Abstract
Automated fake-news detection is increasingly required because the volume and speed of online content exceed the capacity of manual verification. This study presents a controlled comparison between three conventional machine-learning classifiers---Naive Bayes, Logistic Regression, and Random Forest---and a fine-tuned DistilBERT transformer. The conventional models were trained using TF--IDF representations, whereas DistilBERT was fine-tuned directly on tokenized text. Data derived from the LIAR and ISOT fake-news datasets were processed through a common experimental pipeline and evaluated on a held-out test split using accuracy, precision, recall, and F1-score. The reported results show accuracies of 95.67%, 97.93%, and 97.67% for Naive Bayes, Logistic Regression, and Random Forest, respectively, while DistilBERT achieved 99.00%. The trained DistilBERT model was further integrated into a Streamlit application that returns a Real/Fake prediction and an associated confidence score, with a token-level explanation view available in the interface. The study therefore provides a compact benchmark of frequency-based and contextual text representations under a common workflow and demonstrates a lightweight path from model evaluation to interactive deployment.