Skip to content
Review Open access

Balancing Sentiment Analysis Datasets Through Representative-Word-Guided Synthetic Review Generation: A Case Study on Mexican Spanish Tourism Reviews

Jul 2026 · Applied Sciences · 0 citations · 18 references

Abstract

Class imbalance remains one of the most challenging problems in sentiment analysis, particularly in tourism review datasets where positive opinions substantially outnumber neutral and negative comments. This issue is especially critical because minority classes often contain the most valuable information regarding customer dissatisfaction, service failures, and opportunities for improvement. In this work, we propose a hybrid balancing methodology for sentiment analysis in Mexican Spanish tourism reviews that combines undersampling and Large Language Model (LLM)-based oversampling. The proposed framework first extracts representative words from each sentiment class using Mutual Information, then enriches them through dictionary-based or embedding-based lexical substitutions, and finally generates synthetic reviews using GPT-4o-mini guided by these representative terms. Experiments were conducted on a corpus that contains approximately 300,000 tourism reviews collected from TripAdvisor, exhibiting severe sentiment imbalance. Three undersampling strategies and multiple oversampling configurations were evaluated across six traditional machine learning classifiers and one Transformer-based model (BETO). Results show that random undersampling consistently outperformed centroid-based and K-means-based alternatives while also requiring the lowest computational cost. The best overall performance was obtained by BETO, achieving a Macro-F1 score of 0.57 compared to 0.51 on the original imbalanced dataset, representing an improvement of 11.8%. Significant gains were also observed for minority classes, with improvements exceeding 16% for the most underrepresented category. Furthermore, the proposed methodology consistently outperformed direct prompt-based generation using GPT-4o-mini, Gemini 2.5 Flash, and Llama 3.3 70B. These findings suggest that guiding synthetic review generation through representative words effectively preserves domain-specific lexical and semantic patterns of Mexican Spanish tourism reviews, resulting in more balanced datasets and improved sentiment classification performance.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.