Aug 2026· International journal of computer information systems and industrial management applications· Vol 18, pp. 674-687· 0 citations
TL;DR
This study presents a scalable framework for news categorization on a big data platform, improving both effectiveness and efficiency in handling massive datasets and providing an effective NLP-based solution for real-time news classification and intelligent information management in Big Data.
Abstract
With the rapid expansion of digital media, automatic news classification has become essential for managing and filtering the overwhelming volume of online information. Accurate classification enables efficient access to relevant content and supports applications such as personalized recommendations, trend detection, and misinformation control. This study presents a scalable framework for news categorization on a big data platform, improving both effectiveness and efficiency in handling massive datasets. The method integrates Term Frequency–Inverse Document Frequency (TF-IDF) to capture statistical word importance with Bidirectional Encoder Representations from Transformers (BERT) for deep contextual semantics. Hybrid features are processed in a distributed environment using PySpark for fast and reliable computation. By combining TF-IDF and BERT for hybrid feature representation and processing them with an Extra Trees Classifier (ETC) Machine Learning (ML) model, the framework achieves 90.36% accuracy, demonstrating the effectiveness of integrating statistical, contextual, and ensemble-based learning for robust news classification. Experimental evaluation shows the hybrid model outperforms single-model baselines in both accuracy and robustness. By leveraging TF-IDF, BERT, and PySpark, the framework provides an effective NLP-based solution for real-time news classification and intelligent information management in Big Data.
The proposed approach provides a simple and efficient solution for multilingual fake news detection in data-scarce environments with ensemble-based classifiers such as Random Forest and Gradient Boosting achieving reliable performance across both languages.
Nikita Garg, Pritam Singh Negi· International Journal of Eng...· 0 citations
Assamese is a low resource language that presents significant challenges for automated news classification due to the scarcity of curated datasets. To address these challenges, this paper introduces a dedicated corpus of 6,582 Assamese news articles and proposes a hybrid framework that merges a transformer-based subwor...
Pragyat Jyoti Baruah, Arnab Paul, Sourish Dhar et al.· International Conference on...· 0 citations
The exponential growth of textual data on social media and information networks poses a significant challenge to extracting valuable information. Text classification, a core task in Natural Language Processing (NLP), is essential for organizing and categorizing such data. Deep learning has emerged as an effective a...
Ran Jin, Ya Wang, Tianzi Wu et al.· Recent Advances in Computer...· 0 citations
The proposed framework highlights the potential of integrating transformer-based language models with classical machine learning algorithms to build robust and scalable fake news detection systems.
Umme Noor Us Saqa, S. R.· International Journal of Inn...· 0 citations
The rapid growth of user-generated textual content on the internet has intensified the need for accurate and scalable text classification methods. However, supervised learning approaches remain heavily constrained by the high cost and effort required for manual data annotation, particularly in large and heterogeneous d...
A comprehensive natural language processing (NLP) pipeline for extracting key information and identifying industries, which outperforms traditional machine learning baselines and single deep learning models, offering more reliable recognition for minority classes.
Xin-Yi Xu· Applied and Computational En...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.