Skip to content
Open access

Explainable Artificial Intelligence and Computer Vision for RealTime Detection and Prediction of Critical Events in Smart Cities

Aug 2026 · International Journal of Modern Science and Research Technology · 0 citations · 8 references

TL;DR

This study presents XAI-CityVision, an explainable artificial intelligence platform that combines computer vision, object identification, temporal learning, risk assessment, and visual explanation that offers a repeatable architecture for integrating deployment-aware evaluation, operator-oriented explanations, and predictive performance.

Abstract

In order to promote public safety and quick emergency response, smart cities are depending more and more on networked cameras, Internet-of-things sensors, unmanned aerial vehicles, and edge computing. However, occlusion, illumination fluctuation, camera motion, crowd density, complex relationships, and the temporal evolution of anomalous behaviour make it challenging to identify and forecast important events from continuous urban footage. For real-time critical-event monitoring, this study presents XAI-CityVision, an explainable artificial intelligence platform that combines computer vision, object identification, temporal learning, risk assessment, and visual explanation. The framework uses a synchronised acquisition layer to receive heterogeneous urban observations, preprocesses videos, uses CNN or Vision Transformer backbones to extract spatial representations, uses a YOLO-family detector to identify pertinent objects, and uses ConvLSTM or transformer-based temporal learning to model event evolution. Grad-CAM, SHAP, and attention visualisation offer complementary explanations of the choice, while a risk module integrates event likelihood, temporal persistence, and contextual severity. A universal event ontology including normal activity, accident/collision, fire/smoke, crowd anomaly, intrusion/suspicious behaviour, and traffic-related occurrences is mapped to dataset-specific labels in the experimental design, which employs public urban and surveillance benchmarks. Precision, recall, F1-score, IoU, mean average precision, ROCAUC, false-alarm rate, inference delay, and frames per second are used to assess performance. The contributions of spatial learning, object detection, temporal modelling, and multimodal context are measured by ablation studies. In smart-city settings, the suggested framework offers a repeatable architecture for integrating deployment-aware evaluation, operator-oriented explanations, and predictive performance.

Read PDF

Similar papers

Conference Aug 2026

A Hybrid CNN–Vision Language Model Framework for Visual Fire Detection and Context-Aware Decision Support in Oil and Gas Facilities

In Nigeria's oil and gas industry, fire and explosion incidents have claimed more than 3,400 lives over the last 15 years, which are fueled by old infrastructure, chronic pipeline vandalism and lack of intelligent real time monitoring. Traditional sensor-based systems of detection are limited in their spatial coverag...

E. C. Ashinze, E. P. Okoneyo · 0 citations
Conference Open access 2026

A Smart City Framework for Real-Time Traffic Monitoring and Accident Detection

Smart city transport networks must be highly adaptive, meaning they can quickly adjust to new road conditions. This research aims to provide a smart city architecture that can detect accidents and track traffic in realtime using edge-cloud computing, deep learning-based video analytics, and IoT sensing. In order to cor...

R. Elankavi, Imran Alam, Mogadala Mounika et al. · 0 citations
Open access Aug 2026

SmartFire Vision: An Attention-Pruned Hybrid Vision Transformer and Detection Transformer Framework for Accurate, Efficient, and Real-Time Fire and Smoke Detection in Smart City Video Surveillance

Fire incidents can lead to significant destruction of lives and property, especially in urban and smart cities, and pose a great risk worldwide. Existing fire and smoke detection systems are often inadequate for detecting the location of a fire, assessing the speed of its spread, and providing real-time alerts that can...

Muhammad Azhar, Muhammad Arman, Asma Iqbal et al. · 0 citations
Open access 2026

A Deep Spatio-Temporal Framework for Multi-Class Traffic Prediction and Accident Detection in Surveillance Video

Experimental results demonstrate that Z-score standardization improves classification performance, and the feasibility and robustness of the proposed framework in real-world traffic environments are indicated.

Dhartee Patel, Jinal Ahir, Namrata Shroff et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.