Modern Approaches to Visual Content Detection Using YOLO, RF-DETR, and CNN Models
Abstract
This article presents a comparative study of modern deep learning approaches for the detection of destructive visual content, including extremist symbols, weapons, explosive devices, and violent imagery. A domain-specific dataset containing 11 object categories was developed from publicly available Internet sources and manually annotated using bounding boxes in COCO format. The dataset was divided into training (70%), validation (20%), and testing (10%) subsets and expanded through data augmentation to improve model robustness under different viewpoints and background conditions. Several state-of-the-art object detection models were experimentally evaluated, including YOLOv8s, YOLOv8m, YOLOv8l, YOLOv8x, YOLOv11s, an enhanced YOLOv8m model combined with Convolutional Neural Network (CNN) layers, and the transformer-based RF-DETR Medium architecture. Performance was assessed using Precision, Recall, and mean Average Precision at an Intersection over Union threshold of 0.5 (mAP@50). Among the YOLO-based models, YOLOv8m achieved the highest detection accuracy with an mAP@50 of 0.834, whereas the RF-DETR Medium model outperformed the others with an mAP@50 of 0.869, an average Precision of 0.89, and a Recall of 0.85. Class-wise evaluation demonstrated excellent recognition performance for Al-Qaeda, Blood, ISIS, and Hezbollah categories (mAP@50 up to 0.99), while lower accuracy was observed for Hamas and AK-47 because of limited training samples and visual similarity with other categories. The selected RF-DETR Medium model was integrated into a real-time visual threat detection system with a graphical user interface and achieved an average inference speed of approximately 55 FPS, demonstrating its suitability for practical deployment. The obtained results confirm that transformer-based object detectors provide higher detection accuracy and stronger generalization capability than convolution-based approaches for complex destructive visual content, making the proposed system a promising solution for automated content moderation and information security applications.