Skip to content
Open access

A Hybrid Semantic–Sentiment Framework for Automatic Fake News Detection Using Doc2Vec and Machine Learning

Sneha K. Patle Bhushan Gedam Sameer Tembhurney Sachin Balvir A. Agashe Megha Rode Yuvaraj Dakhane Hrushikesh Madhukar Panchabudhe
Aug 2026 · International journal of computer information systems and industrial management applications · 0 citations

TL;DR

An automatic approach for fake news detection which utilizes NLP, sentiment analysis, semantic embedding methods and several machine learning algorithms has been developed in this paper.

Abstract

The spread of misinformation through digital plat-forms such as social media sites and news portals has posed a problem of maintaining information credibility and building trust. Manual approaches alone cannot help cope with the huge volumes of information uploaded on these platforms every day. An automatic approach for fake news detection which utilizes NLP, sentiment analysis, semantic embedding methods and several machine learning algorithms has been developed in this paper. The news headlines collected from FakeNewsNet dataset have been pre-processed via tokenization, stop-word elimination, lemmatization, and n-grams extraction. Doc2Vec approach has been employed to extract semantic vectors whereas sentiment analysis has been done with the help of VADER tool. These semantic vectors have been provided as input to various machine learning algorithms such as Logistic Regression, Linear SVM, Random Forest, Gradient Boosting, XGBoost, LightGBM, Naïve Bayes and K-Nearest Neighbor. Experimental results suggest that ensemble learning models outperform other forms of machine learning techniques. Out of all the tested algorithms, ExtraTrees performed with the highest classification accuracy (77.54%) whereas XGBoost produced the highest macro F1-Score (0.5059).

Read PDF

Similar papers

Review Open access Sep 2026

Feature Based Survey on Fake News Detection: Statistical and Semantic

The digital news portals and social media are rapidly expanding, which has significantly increased the spread of fake news, which affects public opinion, social harmony, and trust in information sources. Detection of fake news at an early stage is a critical research challenge. In recent years, researchers have applied...

Itika U. Lakkewar, R. Jugele · 0 citations
Open access Aug 2026

Fake News Detection Using Machine Learning and LLM Embeddings: A Comparative Study of TF-IDF and BERT Representations on the Welfake Dataset

The proposed framework highlights the potential of integrating transformer-based language models with classical machine learning algorithms to build robust and scalable fake news detection systems.

Umme Noor Us Saqa, S. R. · 0 citations
Open access Aug 2026

Multilingual Fake News Detection Using Machine Learning with Contextual-Based Feature Extraction

The proposed approach provides a simple and efficient solution for multilingual fake news detection in data-scarce environments with ensemble-based classifiers such as Random Forest and Gradient Boosting achieving reliable performance across both languages.

Nikita Garg, Pritam Singh Negi · 0 citations
Open access Jul 2026

Hybrid Lexicon-Driven News Threat Detection Using Random Forest and XGBoost Models

Experimental results indicate that the proposed hybrid approach effectively identifies threatrelated news, with the Random Forest model providing slightly better classification performance than XGBoost.

G. B. Prasad, G.Rajini · 0 citations
Open access Jul 2026

Comparison of Machine Learning Algorithms for Sentiment Analysis of Trans Jogja on Social Media

The experimental results demonstrate that the Support Vector Machine (SVM) consistently outperformed the other algorithms across different data split ratios, indicating that it is the most effective algorithm for sentiment classification of Trans Jogja users on the X platform.

Putri Muryanti Setyowati, Y. Pristyanto, Arif Nur Rohman · 0 citations
Open access Jul 2026

Deep Learning-Based Detection of Machine-Generated Tweets with FastText Word Embeddings

A deep learning approach for detecting machinegenerated tweets using FastText word embeddings and a Convolutional Neural Network and demonstrates better performance than conventional machine learning methods.

Sidhartha.K, Sk.Mahammadunnisa · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.