Development of a Machine Learning-Based Sentiment Analysis Model for Faculty Evaluation in Higher Education
Abstract
Higher education institutions in the Philippines collect faculty evaluation data every semester, yet the qualitative, open-ended comments generated by students remain largely unprocessed and unanalyzed. These narrative responses carry nuanced instructional insights that structured rating scales cannot fully capture. This study developed and evaluated a machine learning-based sentiment analysis model specifically designed to classify student textual feedback on faculty performance into three sentiment categories: positive, neutral, and negative. The dataset comprised 2,500 student-written evaluation comments gathered from a state university in the Philippines, preprocessed through tokenization, stop-word removal, and TF-IDF vectorization, and subsequently used to train and compare five classification algorithms: Naïve Bayes, Support Vector Machine (SVM), Random Forest, Long Short-Term Memory (LSTM), and a fine-tuned Bidirectional Encoder Representations from Transformers (BERT) model. Comparative evaluation on a held-out test set demonstrated that the fine-tuned BERT model delivered the strongest overall classification performance, achieving an accuracy of 92.3%, a precision of 91.8%, a recall of 91.2%, and an F1-score of 91.5%. LSTM followed closely at 89.7% accuracy, while Random Forest, SVM, and Naïve Bayes achieved 85.2%, 83.6%, and 78.4%, respectively. The BERT model’s superior performance is attributed to its deep contextual embeddings and capacity to resolve domain-specific linguistic ambiguity in academic evaluation language. The findings confirm that automated sentiment analysis can substantially augment conventional faculty evaluation systems by surfacing qualitative patterns from student feedback at scale. This study contributes a reproducible NLP pipeline tailored to Philippine academic contexts and lays the groundwork for data-driven instructional quality improvement.