Cyberbullying Detection on Social Media: A Review of Features and Machine Learning Approaches
Abstract
The rapid growth of digital communication platforms has facilitated global connectivity while simultaneously intensifying the prevalence of cyberbullying, a form of online aggression with severe psychological consequences for victims. The research problem addressed in this paper concerns the limitations of manual moderation methods in detecting harmful online behavior at scale. Consequently, this review paper provides a comprehensive synthesis of recent research on cyberbullying detection, with an emphasis on feature engineering and supervised learning algorithms, including Support Vector Machine (SVM) and Random Forest. Feature engineering strategies that include lexical, sentiment, user, and network-based features are examined for their role in enhancing detection accuracy. Analysis of 13 key studies reveals that SVM and Random Forest consistently outperform other classifiers, especially when combined with robust optimization techniques. The review identifies that content-based features, particularly TF-IDF, combined with hyperparameter tuning, achieve accuracies exceeding 94% in optimal configurations. The study concludes that while current models demonstrate substantial progress, challenges related to linguistic ambiguity, data imbalance, cross-platform generalizability, and lack of multilingual support remain critical directions for future research.