Skip to content
Open access

Modern Approaches to Visual Content Detection Using YOLO, RF-DETR, and CNN Models

Oct 2026 · Engineering, Technology & Applied Science Research · 0 citations · 19 references

Abstract

This article presents a comparative study of modern deep learning approaches for the detection of destructive visual content, including extremist symbols, weapons, explosive devices, and violent imagery. A domain-specific dataset containing 11 object categories was developed from publicly available Internet sources and manually annotated using bounding boxes in COCO format. The dataset was divided into training (70%), validation (20%), and testing (10%) subsets and expanded through data augmentation to improve model robustness under different viewpoints and background conditions. Several state-of-the-art object detection models were experimentally evaluated, including YOLOv8s, YOLOv8m, YOLOv8l, YOLOv8x, YOLOv11s, an enhanced YOLOv8m model combined with Convolutional Neural Network (CNN) layers, and the transformer-based RF-DETR Medium architecture. Performance was assessed using Precision, Recall, and mean Average Precision at an Intersection over Union threshold of 0.5 (mAP@50). Among the YOLO-based models, YOLOv8m achieved the highest detection accuracy with an mAP@50 of 0.834, whereas the RF-DETR Medium model outperformed the others with an mAP@50 of 0.869, an average Precision of 0.89, and a Recall of 0.85. Class-wise evaluation demonstrated excellent recognition performance for Al-Qaeda, Blood, ISIS, and Hezbollah categories (mAP@50 up to 0.99), while lower accuracy was observed for Hamas and AK-47 because of limited training samples and visual similarity with other categories. The selected RF-DETR Medium model was integrated into a real-time visual threat detection system with a graphical user interface and achieved an average inference speed of approximately 55 FPS, demonstrating its suitability for practical deployment. The obtained results confirm that transformer-based object detectors provide higher detection accuracy and stronger generalization capability than convolution-based approaches for complex destructive visual content, making the proposed system a promising solution for automated content moderation and information security applications.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.