Deep Learning and Edge Computing-Based Weapon Detection Models for Enhancing Perimeter Security in Smart Cities
Abstract
Early firearm detection in urban environments represents an important challenge for modern security and video surveillance systems. In this context, this work proposes and evaluates detection models based on Deep Learning, TinyML, and Edge AI for embedded devices and edge-computing platforms. The research integrates architectures such as MobileNetV2, YOLOv5, YOLOv7, YOLOv8, and Vision-Language Models (VLMs) to analyze their feasibility for real-time video surveillance applications. The experimental methodology was developed in three main scenarios. First, a TinyML-based system using an ESP32-S3 microcontroller and the FOMO-MobileNetV2 model was implemented to evaluate local inference on low-power hardware. Subsequently, TensorRT-optimized YOLO models were deployed on an NVIDIA Jetson Nano platform to compare inference speed and real-time processing capabilities. Finally, Vision-Language Models supported by MediaPipe were incorporated to improve contextual scene analysis and reduce computational workload through dynamic regions of interest. The results demonstrate that TinyML and Edge AI enable artificial intelligence models to run directly on edge devices, reducing latency and increasing the operational autonomy of surveillance systems. In addition, YOLO architectures showed high performance in real-time object detection, while Vision-Language Models provided enhanced contextual understanding. Overall, the findings confirm the feasibility of developing intelligent, scalable, and low-cost solutions for perimeter security in smart cities.