Deep Learning and Computer Vision-Based Modeling for a Real-Time Traffic Congestion Monitoring Framework
Abstract
As intelligent traffic systems evolve to manage complex urban mobility, conventional congestion estimation techniques, such as the time-windowed Volume-to-Capacity (V/C) ratio, fail to capture capture the real-time traffic situation. Because these methods rely on vehicles crossing a specific point, they often fail to register stopped cars, creating a ‘zero flow’ paradox where complete gridlock is incorrectly categorized as an empty, free-flowing road. To address these limitations, this study proposes an end-to-end computer vision framework utilizing the Real-Time DEtection TRansformer (RT-DETR) model to quantify vehicle volume and immediately classify a roadway’s Level of Service (LOS) through spatial density. In processing video feeds from the Ateneo de Manila University Campus Safety and Mobility Office, the framework calculates an instantaneous, frame-level Spatial Density Index across six distinct vehicle classes. Results show that the proposed SDI framework outperforms the traditional time-windowed method in estimating traffic congestion and reduces percentage error margins to between -0.22% and 2.85%. The proposed framework successfully resolves the zero flow paradox by analyzing frame-wide congestion rather than fixed line crossings prescribed in the time-windowed method. It also offers a highly responsive and robust alternative that detects immediate spikes in vehicle density, which makes it potentially suitable for integration into intelligent traffic systems.