ViTBN: A Vision Transformer with Batch Normalization for Crowd Behaviour and Anomaly Detection in Surveillance Videos
Abstract
Intelligent crowd behaviour analysis is critical for modern video surveillance systems to enhance public safety and enable timely anomaly detection. This study presents ViTBN, a vision transformer-based framework augmented with batch normalization to improve feature robustness and the stability of training. The proposed framework processes frames extracted from surveillance videos and employs a pretrained Vision Transformer to capture long-range spatial relationships within individual frames. The extracted representations are subsequently processed through feature aggregation, Batch normalization, and a classification head for normal anomalous discrimination. Experimental evaluation on the UCF-Crime benchmark demonstrates an accuracy of 91.29%, precision of 83.46%, recall of 87.23%, and F1-score of 85.33%, respectively, indicating reliable discrimination between anomalous and normal surveillance activities. Ablation studies confirmed that incorporating batch normalization improved optimization stability and feature generalization compared to the baseline ViT setup. The findings suggest that integrating normalization layers into transformer-based architectures strengthens their applicability in surveillance contexts, offering a practical and effective approach to intelligent public safety systems in real-world settings