DeepVisionSafety: An Intelligent CCTV Monitoring System for Multi-Class Anomaly Recognition
Abstract
This study presents a supervised learning framework for abnormal event recognition in CCTV footage. The proposed architecture utilizes the Inflated 3D ConvNet (I3D) for the video feature extraction task, followed by a lightweight classification head consisting of fully connected layers, batch normalization, ReLU activation, and dropout regularization. To evaluate the model’s efficiency, we employed multiple evaluation metrics, including Accuracy and Area Under the Curve (AUC). The experiments were benchmarked using the UCF-Crime and Gun Action Recognition datasets, alongside a custom-curated dataset. Experimental results indicate that the proposed I3D (RGB) model achieves a multi-class accuracy of 88.0% and a binary AUC of 0.981 with a feature extraction time of 2.11 seconds, demonstrating superior performance in controlled indoor environments, such as rooms or building interiors.