PSO-Based Hybrid Lexical-Semantic Feature Selection and Ensemble Learning Framework for Urdu Hate Speech Detection
Abstract
Automated Urdu hate-speech detection remains challenging because annotated resources are limited, orthography varies, and harmful meaning depends on both lexical and contextual cues. This study evaluates a leakage-safe lexical-semantic framework combining 5,000-dimensional word/bigram TF-IDF features with 768-dimensional multilingual sentence embeddings, followed by binary particle swarm optimisation (BPSO) and classical ensemble learning. Experiments used the Urdu text and binary labels from MMHS11K: 8,800 balanced training records and an untouched balanced test set of 2,200 records. BPSO selected 2,849 of 5,768 hybrid dimensions, reducing dimensionality by 50.61%. On the official test set, Stacking with the complete hybrid representation achieved Macro-F1 = 0.8468 and ROC-AUC = 0.9265; the PSO-selected representation achieved Macro-F1 = 0.8391 and ROC- AUC = 0.9196. After mask selection, classifier fitting time decreased by 52.24%, excluding sentence-embedding extraction and BPSO search. Leakage-safe nested five-fold validation showed a similar trade-off: fitting time decreased by 52.64%, with Macro-F1 = 0.8444 ± 0.0066 after PSO versus 0.8492 ± 0.0110 without PSO. Soft Voting and Stacking were statistically comparable on selected features. Three-seed transformer aggregates performed better: Urdu-RoBERTa achieved the highest Macro-F1 (0.8836), and XLM-R achieved the highest ROC-AUC (0.9529). Holm-corrected exact McNemar tests confirmed that transformer aggregates significantly outperformed PSO Stacking. Overall, BPSO approximately halves representation size and downstream classical-model fitting time, with a small predictive decrease that was not statistically significant relative to complete-hybrid Stacking after correction.