Skip to content
Review Open access

Popularity Bias and Category Diversity in Review-Based Rating Prediction for E-Commerce: An Empirical Study on Indonesian Marketplace Data

2026 · International Journal of Advanced Computer Science and Applications · 0 citations · 37 references

Abstract

Online marketplaces increasingly rely on user reviews to estimate product quality. However, most empirical comparisons of review-based prediction models report only aggregate accuracy and overlook how item popularity and category composition shape that accuracy. This study presents a controlled empirical analysis of review-based rating prediction on a self-collected, anonymized corpus of 44,229 Indonesian marketplace reviews spanning four product categories (fashion, electronics, tools/hardware, and sports). To isolate the effect of category diversity from data volume, the experimental design was set with the total number of reviews kept constant at 12,000, while the number of categories was increased from two to four. For product metadata classification, three text classification models (TF-IDF with logistic regression, linear SVM, and random forest) and a customized IndoBERT transformer model were compared using five-fold cross-validation, and the difficulty of prediction was further analyzed at the item level. Three findings emerge. First, the review text is the dominant signal, reducing the mean absolute error to 0.62 (0.57 with IndoBERT), compared with 1.28 for the majority baseline. Second, increasing category diversity within a fixed budget results in a small but consistent performance degradation. Third, and most notably, popular items are systematically harder to predict than long-tail items within every category, and per-item error correlates positively with rating variance and sales volume and negatively with average product rating and price. The polarization of evaluations is considered a major contributing factor to production difficulties; the protocol can be reproduced and interpreted at a level where evaluation predictions are based on a pass/fail assessment.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.