Skip to content
Open access

Comparison of SMOTE, Class Weighting, and Classical Machine Learning Models on the ID-SMSA Indonesian Stock Market Dataset

Jul 2026 · Journal of System and Computer Engineering (JSCE) · Vol 7, pp. 229-240 · 0 citations · 24 references

TL;DR

A transparent and reproducible classical baseline is provided that situates transformer-based and deep-learning approaches on ID-SMSA within a well-defined reference frame and confirms that linear SVM on TF-IDF is robust to moderate imbalance at IR = 2.41.

Abstract

Sentiment classification of social-media text related to the Indonesian stock market is a growing research area. The ID-SMSA dataset is the publicly available labelled corpus for this domain, yet class-imbalance handling strategies on this dataset have not been systematically compared across multiple classifiers. This paper evaluates Multinomial Naive Bayes, linear Support Vector Machine (SVM), and Random Forest under three imbalance-handling conditions: no handling, class weighting, and SMOTE. All experiments use the full 3,287-tweet dataset with an 80:20 stratified split and report macro F1 as the primary metric. SMOTE consistently improves macro F1 across all classifiers. The largest gain is on Naive Bayes (+0.137, from 0.589 to 0.726). The best configuration is SVM with SMOTE, achieving macro F1 of 0.752 and accuracy of 0.784. Class weighting benefits Random Forest (+0.011) but slightly reduces SVM, confirming that linear SVM on TF-IDF is robust to moderate imbalance at IR = 2.41. Per-issuer evaluation reveals macro F1 variation from 0.647 on TPIA to 0.881 on BBNI, shaped by vocabulary consistency, class dominance, and domain specificity. These results provide a transparent and reproducible classical baseline that situates transformer-based and deep-learning approaches on ID-SMSA within a well-defined reference frame.

Read PDF

Similar papers

Review Open access Sep 2026

Performance of Classical Machine Learning Algorithms in Three-Class Sentiment Classification of Indonesian DANA E-Wallet Reviews

Automated sentiment classification of e-wallet reviews can support service monitoring, yet reported performance may be inflated by duplicate texts, conflicting labels, and test-set-driven model selection. This study compares Multinomial Naive Bayes (MNB), linear Support Vector Machine (SVM), and Logistic Regression (LR...

Husni Mubarak, Clara Diva, La Ode Fefli Yarlin et al. · 0 citations
Review Open access Sep 2026

Machine Learning-Based Sentiment Classification of Reviews from Indonesian Mobile Applications Using TF-IDF

Comparisons of classical machine learning algorithms and the effectiveness of class-weighted learning in improving minority-class recognition in Indonesian mobile application reviews demonstrate that the highest overall accuracy does not necessarily indicate the most balanced classifier under class imbalance.

Tuti Handayani, Sri Mardiyati · 0 citations
Review Open access Sep 2026

Improving Sentiment Classification Performance Using Pseudo-Labeling with Naive Bayes and Random Forest

Overall, TF-IDF outperformed Count Vectorizer, and larger threshold values yielded more consistent performance improvements across datasets, though lower values offered greater potential for gains on large, diverse datasets, which suggest pseudo-labeling is a viable method for incorporating unlabeled data.

Arvidion Havas Oktavian, A. Aribowo · 0 citations
Open access Aug 2026

SMOTE-Stack-XAI: An Explainable Stacked Ensemble Learning Framework Integrating Random Forest, XGBoost, SVM and Deep Neural Networks for Real-Time Credit Card Fraud Detection

SMOTE-Stack-XAI, a stacked ensemble framework that uses the Synthetic Minority Oversampling Technique (SMOTE) to address class imbalance, outperforms the RF, hybrid RF-SVM, and hybrid SVM-LR baselines with an accuracy of 99.3%, precision of 99.2%, recall of 99.4%, and F1-score of 99.3%.

Bhukya Dharma, D. Latha · 0 citations
Review Sep 2026

EVALUATION OF TEXT CLASSIFICATION ALGORITHMS: A STUDY ON CLASS IMBALANCE AND PREPROCESSING

This paper presents a controlled comparative study of traditional machine learning algorithms for thematic text classification, focusing on the impact of preprocessing strategies and class imbalance on model performance under different data conditions. Two experimental scenarios were considered: the Women’s Clothing E-...

Dulce Liliana Estrada Bahena, Alicia Martínez Rebollar, Hugo Estrada Esquivel et al. · 0 citations
Open access Sep 2026

Comparison of Naive Bayes, SVM, and Logistic Regression for Sentiment Analysis of the Makan Bergizi Gratis Program

The Makan Bergizi Gratis Program (MBG) became one of the widely discussed public issues on platform X and generated diverse responses from users. These responses included supportive, critical, and neutral opinions, making sentiment analysis relevant for understanding public opinion toward the program. This study compar...

Andika, Julio Cesar Alessandro, Adjie Perkasa Tarigan et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.