Sentiment Analysis of Imbalanced Dataset Through Data Augmentation and Generative Annotation Using DistilBERT and Low‐Rank Fine‐Tuning
Abstract
Sentiment analysis on social media data often suffers from severe class imbalance, which can negatively affect the performance of machine learning models. In this paper, we propose a framework that leverages large language models and lightweight transformer fine tuning to improve sentiment classification on imbalanced datasets. First, GPT‐4, a multimodal large language model, is used to generate synthetic tweets through paraphrasing and back translation, with Italian serving as an intermediate language, to increase data diversity. In addition, GPT‐4 is employed to annotate tweets with positive reasons by generating semantic counterparts to the 10 predefined negative categories in the Twitter US Airline Sentiment dataset. This process enables the creation of meaningful positive annotations derived from existing category structures, thereby improving dataset balance and interpretability. The augmented data are then encoded using DistilBERT to obtain sentence embeddings, while low rank adaptation (LoRA) is applied for efficient fine tuning with reduced computational cost. Finally, a SoftMax classifier is used to predict sentiment labels (positive, neutral, and negative). Experimental results on the Twitter US Airline Sentiment dataset—evaluated rigorously using a held‐out test set and 10‐fold cross‐validation—demonstrate that the proposed framework achieves strong classification performance while maintaining low training complexity, highlighting the effectiveness of combining Large Language Model (LLM) based data augmentation with parameter‐efficient transformer adaptation.