Skip to content
Open access

News Classification Using Hybrid Natural Language Processing Based Content Modelling with Machine Learning Methods

Aug 2026 · International journal of computer information systems and industrial management applications · Vol 18, pp. 674-687 · 0 citations

TL;DR

This study presents a scalable framework for news categorization on a big data platform, improving both effectiveness and efficiency in handling massive datasets and providing an effective NLP-based solution for real-time news classification and intelligent information management in Big Data.

Abstract

With the rapid expansion of digital media, automatic news classification has become essential for managing and filtering the overwhelming volume of online information. Accurate classification enables efficient access to relevant content and supports applications such as personalized recommendations, trend detection, and misinformation control. This study presents a scalable framework for news categorization on a big data platform, improving both effectiveness and efficiency in handling massive datasets. The method integrates Term Frequency–Inverse Document Frequency (TF-IDF) to capture statistical word importance with Bidirectional Encoder Representations from Transformers (BERT) for deep contextual semantics. Hybrid features are processed in a distributed environment using PySpark for fast and reliable computation. By combining TF-IDF and BERT for hybrid feature representation and processing them with an Extra Trees Classifier (ETC) Machine Learning (ML) model, the framework achieves 90.36% accuracy, demonstrating the effectiveness of integrating statistical, contextual, and ensemble-based learning for robust news classification. Experimental evaluation shows the hybrid model outperforms single-model baselines in both accuracy and robustness. By leveraging TF-IDF, BERT, and PySpark, the framework provides an effective NLP-based solution for real-time news classification and intelligent information management in Big Data.

Read PDF

Similar papers

Open access Aug 2026

Multilingual Fake News Detection Using Machine Learning with Contextual-Based Feature Extraction

The proposed approach provides a simple and efficient solution for multilingual fake news detection in data-scarce environments with ensemble-based classifiers such as Random Forest and Gradient Boosting achieving reliable performance across both languages.

Nikita Garg, Pritam Singh Negi · 0 citations
Conference Aug 2026

A Hybrid Framework for Automated News Classification for the Low-Resource Assamese Language

Assamese is a low resource language that presents significant challenges for automated news classification due to the scarcity of curated datasets. To address these challenges, this paper introduces a dedicated corpus of 6,582 Assamese news articles and proposes a hybrid framework that merges a transformer-based subwor...

Pragyat Jyoti Baruah, Arnab Paul, Sourish Dhar et al. · 0 citations
Review Aug 2026

A Review of Deep Learning-Based Text Classification Research

The exponential growth of textual data on social media and information networks poses a significant challenge to extracting valuable information. Text classification, a core task in Natural Language Processing (NLP), is essential for organizing and categorizing such data. Deep learning has emerged as an effective a...

Ran Jin, Ya Wang, Tianzi Wu et al. · 0 citations
Open access Aug 2026

Fake News Detection Using Machine Learning and LLM Embeddings: A Comparative Study of TF-IDF and BERT Representations on the Welfake Dataset

The proposed framework highlights the potential of integrating transformer-based language models with classical machine learning algorithms to build robust and scalable fake news detection systems.

Umme Noor Us Saqa, S. R. · 0 citations
Open access Aug 2026

A novel hybrid model for identifying the most informative instances for improving text data classification

The rapid growth of user-generated textual content on the internet has intensified the need for accurate and scalable text classification methods. However, supervised learning approaches remain heavily constrained by the high cost and effort required for manual data annotation, particularly in large and heterogeneous d...

A. Abdelwahab, M. Salama · 0 citations
#large language models Open access Sep 2026

Research on Text Information Extraction and Imbalanced Classification Methods for Enterprise Profiling

A comprehensive natural language processing (NLP) pipeline for extracting key information and identifying industries, which outperforms traditional machine learning baselines and single deep learning models, offering more reliable recognition for minority classes.

Xin-Yi Xu · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.