Jul 2026· Jurnal Informatika· pp. 335-341· 0 citations
TL;DR
It is demonstrated that semantic representations and controlled lexical variation can jointly enhance minority-class recognition in short informal Indonesian text and highlight the importance of aligning embedding strategies with sequence architectures.
Abstract
The rapid expansion of digital payments has produced massive volumes of user-generated reviews, making manual analysis impractical. This study focuses on the challenge of neutral sentiment classification in Indonesian e-wallet reviews, where neutral comments often contain ambiguous language and are underrepresented relative to positive and negative classes. A total of 26,537 preprocessed DANA application reviews were used to evaluate whether Word2Vec embeddings and Easy Data Augmentation (EDA) can improve neutral sentiment detection when combined with Long Short-Term Memory (LSTM) and Bidirectional Long Short-Term Memory (BiLSTM) architectures. Experiments comparing eight model configurations showed that the combination of Word2Vec, EDA, and LSTM achieved the best performance, with 0.861 accuracy, 0.841 macro-F1, and 0.749 F1-score for the neutral class. These findings demonstrate that semantic representations and controlled lexical variation can jointly enhance minority-class recognition in short informal Indonesian text and highlight the importance of aligning embedding strategies with sequence architectures.
A Hybrid VADER–IndoBERT framework designed to improve sentiment classification robustness on complex Indonesian texts is introduced, demonstrating the superiority of Transformer-based architectures in capturing long-range dependencies and handling ambiguous sentiment cues.
Margareta Valencia Suci Handayani, R. S. Basuki, Muljono et al.· Jurnal RESTI (Rekayasa Siste...· 0 citations
This study aims to analyze public sentiment toward the LPDP alumni controversy on social media using a deep learning approach. The research data consist of YouTube user comments related to the LPDP issue, which were processed through text preprocessing and automatically labeled using IndoBERT into three sentiment classes: negative, neutral, and positive. This study compares two text representation methods, namely Word2Vec and FastText, implemented within a hybrid CNN–BiLSTM architecture. In addition, data imbalance was addressed using class weighting and undersampling scenarios, while TF-IDF-based Logistic Regression was used as the baseline model. The results show that the baseline achieved an accuracy of 0.83 but was strongly biased toward the negative class as the majority class. The CNN–BiLSTM model improved the ability to detect minority classes. Under the class weighting scenario, FastText demonstrated more stable performance with an accuracy of 0.77 and a macro F1-score of 0.62. Under the undersampling scenario, Word2Vec was more stable, achieving an accuracy of 0.68 and a macro F1-score of 0.67. These findings indicate that both text representation and imbalance-handling strategies substantially affect sentiment classification performance.
Dwi Erzalianti, Joice Junansi Tandirerung, C. Suhaeni et al.· JOURNAL OF APPLIED INFORMATI...· 0 citations
Sentiment analysis has become an important task in natural language processing for understanding public opinions expressed in online reviews. However, most publicly available IMDb datasets are limited to binary sentiment labels, which restricts the ability of sentiment analysis systems to capture neutral opinions. This study proposes an efficient sentiment analysis framework that transforms the binary IMDb dataset into a three-class sentiment classification problem consisting of positive, neutral, and negative sentiments. The proposed approach integrates pseudolabeling with Parameter-Efficient Fine-Tuning (PEFT) using the Low-Rank Adaptation (LoRA) technique on the Longformer architecture. Experimental results show that the model achieves an accuracy of 77.06%, a weighted F1-score of 72.17%, and a Matthews Correlation Coefficient (MCC) of 0.6232. The results demonstrate that LoRA-based fine-tuning can significantly reduce computational requirements while maintaining competitive performance in sentiment classification tasks. These findings indicate that the proposed framework provides a practical and computationally efficient solution for large-scale sentiment analysis, particularly for environments with limited computational resources.
P. Hiskiawan, Wendy Tjung, Dustin Darmawan Isya Widjaja et al.· JRST: Jurnal Riset Sains dan...· 0 citations
Indonesia's expanding e-commerce sector generates a growing volume of customer-written product reviews that can reveal both satisfaction and dissatisfaction. Automatically determining sentiment in these reviews is nevertheless difficult because marketplace language commonly includes informal wording, inconsistent spelling, brief statements, and domain-specific terms. This research benchmarks conventional machine learning methods for classifying the sentiment of Indonesian e-commerce reviews in the PRDECT-ID dataset. The data were obtained from Tokopedia and contain sentiment and emotion annotations. Following preprocessing, the experiment used 5,305 reviews, comprising 2,752 negative and 2,553 positive instances. The processing pipeline included case folding, text cleaning, normalization, tokenization, selective removal of stopwords, and Term Frequency-Inverse Document Frequency (TF-IDF) feature construction. Multinomial Naive Bayes, Support Vector Machine, and Random Forest were then evaluated under the same experimental configuration. The TF-IDF and Support Vector Machine combination produced the strongest results, reaching 0.9595 accuracy, 0.9594 macro-F1, and 0.9595 weighted-F1. These findings establish a reproducible reference point for sentiment classification in Indonesian e-commerce reviews.
Muhammad Fairuzabadi, Indo Intan, Sitti Suhada· JTH: Journal of Technology a...· 0 citations
Nowadays, Natural Language Processing, or NLP, is a key component of many programs that analyze and comprehend human language. The sentiment analysis of mobile product reviews collected from the Kaggle repository—more especially, the 20,710-review Amazon Mobile evaluations dataset—is the main emphasis of this research. Reviews of well-known cellphone companies including Samsung, Nokia, Apple, Redmi, and others are included in the dataset. The main goal of this research is to categorize consumer attitudes into three groups: neutral, negative, and positive. This research heavily relies on Natural Language Processing (NLP), particularly in the commercial and e-commerce domains where decision-making and product enhancement depend on a comprehension of client input. In this research, sentiment categorization is carried out using deep learning methods like Long Short-Term Memory (LSTM), Artificial Neural Network (ANN), and Convolutional Neural Network (CNN). Lowercase conversion, punctuation and symbol removal, tokenization, stopword removal, stemming, and lemmatization are some of the preprocessing methods used to enhance text quality and model performance. The algorithms are compared based on execution time and accuracy. The survey also determines the top-performing mobile brand by counting the amount of positive reviews. The outcomes show that deep learning models perform better and produce the best results. This research promotes business intelligence in the e-commerce industry and advances our understanding of consumer sentiment behavior.
Dhananchezhiyan R, M. Rameshkumar· International journal of com...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.