Skip to content
Open access

Multi-level violence recognition via hybrid convolutional-attention and recurrent architectures

Aug 2026 · Scientific Reports · 0 citations

Abstract

Violence recognition is an urgent need for real-time human activity identification in surveillance video streams, which is becoming more important in areas including public safety, law enforcement, and security monitoring. Even while visual understanding has come a long way, finding a balance between accuracy and speed of identification is still a big problem. To address these constraints, we suggest a hybrid deep learning architecture that combines a Time-Distributed Convolutional Neural Network (TD-CNN) with a Long Short-Term Memory (LSTM) network augmented by a spatial attention mechanism. The attention module, which is located between convolutional layers, adaptively highlights important spatial areas, which makes it easier to tell the features apart. The LSTM part of the model looks at how frames are related to each other across time to capture motion dynamics that are important to violent behaviors. The suggested architecture was tested on the Hockey Fight Dataset and got an accuracy of 93%, which is better than many other baselines. Experimental findings indicate that the hierarchical integration of convolutional and recurrent layers with attention significantly improves recognition accuracy while maintaining a computationally lightweight architecture suitable for near-real-time deployment.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.