Efficient video violence detection through segmented temporal sampling and CNN–transformer modeling
Detecting violence in video requires models capable of capturing spatial and temporal information without relying on long sequences, optical flow, or additional modalities. This study presents a CNN–Transformer framework based on segmented temporal sampling, evaluating 4, 8, and 16 frames per video. Each frame is encod...