Indonesian Hate Speech Detection: A Cross-Validated Benchmark of Machine Learning and Pre-trained Transformer Models with Statistical Significance Analysis
Automated hate speech detection in Indonesian social media remains a persistent challenge due to dataset fragmentation, heterogeneous annotation schemes, and the lack of reproducible cross-model benchmarks with formal statistical validation. This study presents a cross-validated benchmark that systematically evaluates...