Efficient video violence detection through segmented temporal sampling and CNN–transformer modeling
Abstract
Detecting violence in video requires models capable of capturing spatial and temporal information without relying on long sequences, optical flow, or additional modalities. This study presents a CNN–Transformer framework based on segmented temporal sampling, evaluating 4, 8, and 16 frames per video. Each frame is encoded using EfficientNet-B0 pre-trained on ImageNet; features are projected onto 256-dimensional tokens, supplemented with sinusoidal positional encoding, and processed using a Transformer encoder followed by attention pooling. The framework was evaluated using video-level partitioning, multi-seed training, bootstrapping, robustness analysis, computational efficiency, independent benchmark evaluation and cross-dataset generalization. The 4-frame configuration achieved the most favorable multi-seed validation performance and was therefore selected for final evaluation. On the RLVS test set, the selected configuration achieved an Accuracy of 0.9508 ± 0.0038, F1-score of 0.9517 ± 0.0039, MCC of 0.9022 ± 0.0078, and PR-AUC of 0.9915 ± 0.0008. Under Gaussian noise, the CNN–Transformer showed the smallest MCC degradation (ΔMCC = −0.1853). Independent evaluation yielded Accuracies of 0.9517 ± 0.0076 on Hockey Fight Videos and 0.8857 ± 0.0143 on AIRTLab. The final model contained 5.39 million parameters and achieved 9.57 ± 1.25 ms inference latency per video. These results support segmented temporal sampling as an efficient strategy for violence detection while highlighting remaining challenges under domain shift.