Jul 2026· Jurnal ilmu komputer dan aplikasi· Vol 9, pp. 39-49· 1 citation· 18 references
TL;DR
Comparing three imbalance-handling strategies showed that cost-sensitive learning (class-weight) was most effective, raising the neutral-class F1-score from 26.56% to 35.8% without compromising majority-class performance.
Abstract
The rapid growth of telemedicine services in Indonesia has generated a substantial volume of user reviews, an important source for evaluating service quality. This study compares the effectiveness of lexicon-based and machine learning approaches in classifying review sentiment, using the three largest telemedicine applications in Indonesia, Alodokter, Halodoc, and KlikDokter as a case study. A publicly available dataset of 19,389 reviews, manually labelled by two annotators under the guidance of a psychologist, served as the gold standard and was preprocessed through case folding, normalisation, stopword removal, and stemming. The lexicon-based approach employs the InSet dictionary, while the machine learning approach applies Support Vector Machine (SVM), Naïve Bayes (NB), and Random Forest (RF) with TF-IDF features, evaluated via stratified 10-fold cross-validation. Given the highly imbalanced class distribution (positive 74.77%, negative 21.69%, neutral 3.54%), macro-F1 is adopted as the primary metric alongside accuracy. Machine learning approaches substantially outperformed the lexicon-based approach: SVM achieved the highest macro-F1 of 70.43% (88.68% accuracy), far exceeding the InSet dictionary's macro-F1 of 27.67% (34.92% accuracy). A further key finding is an evaluation paradox: although Naïve Bayes attained the highest accuracy (89.91%), its macro-F1 was comparatively low (60.11%) due to bias toward the majority class, demonstrating how accuracy alone can mislead on imbalanced data. The InSet dictionary also performed poorly in the telemedicine domain, frequently misclassifying positive reviews as negative. Finally, comparing three imbalance-handling strategies (no handling, class-weight, and random oversampling) showed that cost-sensitive learning (class-weight) was most effective, raising the neutral-class F1-score from 26.56% to 35.8% without compromising majority-class performance.
Comparisons of classical machine learning algorithms and the effectiveness of class-weighted learning in improving minority-class recognition in Indonesian mobile application reviews demonstrate that the highest overall accuracy does not necessarily indicate the most balanced classifier under class imbalance.
Tuti Handayani, Sri Mardiyati· Journal Mobile Technologies...· 0 citations
This study develops an automated sentiment analysis system for classifying reviews of the RSI Sunan Kudus mobile application, collected from Google Play Store, Google Maps, and YouTube (N = 1,428). Sentiment labels were automatically assigned using a domain-adapted Indonesian lexicon (87 positive / 100 negative terms),...
Fikri Hamdhan Dwi Saputra, Fajar Nugraha, Zainur Romadhon· INOVTEK Polbeng - Seri Infor...· 0 citations
Mobile application marketplaces generate huge volumes of user reviews daily. Automatically understanding if a review is positive or negative matters to developers, platform operators, and prospective users alike, since manually reading them one by one is not feasible. This paper reviews the empirical and methodologica...
C. E. Orie, A. Egwali, Frank Iwebuke Amadin· Journal of Science Research...· 0 citations
It is demonstrated that the TF–IDF-weighted Naïve Bayes pipeline effectively classifies Indonesian-language health service reviews and yields evidence-based priorities for improving the Mobile JKN platform.
Sigit Andriyanto, Hardo Ardiyanto, T. Nugroho et al.· 0 citations
The rapid growth of the skincare industry has generated a massive volume of consumer reviews on e-commerce platforms, making manual sentiment analysis increasingly impractical. This study compares the performance of Multinomial Naïve Bayes and Support Vector Machine (SVM) for sentiment classification of skincare produc...
R. Putri, Reykha Putri Randika, Febi Dwi Sasmita et al.· JOURNAL OF APPLIED INFORMATI...· 0 citations
Overall, TF-IDF outperformed Count Vectorizer, and larger threshold values yielded more consistent performance improvements across datasets, though lower values offered greater potential for gains on large, diverse datasets, which suggest pseudo-labeling is a viable method for incorporating unlabeled data.
Arvidion Havas Oktavian, A. Aribowo· TEPIAN· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.