Skip to content
Open access

Safety-Aware Reinforcement Learning Model for Adaptive Traffic Signal Optimization in Work Zone Environments

Aug 2026 · Future Transportation · 1 citation · 86 references

TL;DR

The findings show that a single controller trained with surrogate safety indicators as learning objectives can improve operational performance while reducing safety-critical instability in work zones.

Abstract

Work zones reduce roadway capacity and create unstable merging, queue spillback, and stop-and-go conditions that degrade traffic operations while elevating crash risk. Conventional fixed-time, actuated, and adaptive controllers are poorly suited to these non-stationary conditions, and most reinforcement learning approaches optimize mobility while treating safety only as a post hoc evaluation measure. This study develops a safety-aware Deep Q-Network framework for adaptive signal control at intersections operating near work zone activity areas. Merge conflict risk, upstream spillback propagation, and stop-and-go instability are embedded directly into both the state representation and the reward formulation, alongside operational objectives. A merge-conflict model based on relative spacing, relative speed, and acceleration characterizes unsafe interactions in the merge region, and a Pareto-based procedure samples reward-weight vectors to identify non-dominated policies. The framework was evaluated in a SUMO microscopic simulation of a signalized intersection under lane closure. Relative to default fixed-time control, the selected policy increased throughput by 24.6–37.3% across vehicle classes (p < 0.001; Cohen’s d = 0.53–1.29), with the largest gains for trucks and buses, and reduced maximum queue length by 39.1% and spillback distance by 45.8%. The findings show that a single controller trained with surrogate safety indicators as learning objectives can improve operational performance while reducing safety-critical instability in work zones.

Read PDF

Similar papers

Open access Aug 2026

Deep reinforcement learning-based traffic signal control in multi-intersection environments: a comparative study of DQN variants

The findings demonstrate the potential of DRL-based traffic signal control in controlled simulation conditions and highlight that algorithm performance is strongly influenced by traffic policy design and environmental complexity.

D. Prastiyanto, A. A. Manaf, Muhammad Ahnaf Maulana et al. · 0 citations
Open access Jul 2026

Traffic Signal Control via Proximal Policy Optimization with Reward Shaping to Minimize Waiting Time and Violations

Traffic signal control at intersections is a key challenge in urban traffic management, particularly when accounting for non-compliant driver behavior such as red-light violations. While Proximal Policy Optimization (PPO) has shown promise for adaptive traffic control, most implementations overlook the stochastic nature of such violations, limiting real-world applicability. This study proposes an enhanced PPO-based traffic signal control approach that incorporates reward shaping to minimize vehicle waiting time and traffic violations jointly. The method modifies the reward function by adding a fixed bonus when no violations occur and a logarithmically scaled penalty when violations are detected. Experiments were conducted in Simulation of Urban Mobility (SUMO) using a real-world intersection model, with aggressive and violation-prone driver behavior generated through domain randomization. Evaluation covered three training scenarios (5%, 10%, and 20% violation rates) and two additional test scenarios with different intersection layouts. In training scenarios, PPO with reward shaping reduced violations to 3-4 while maintaining moderate delays of 14-22 seconds, outperforming PPO without reward shaping, which primarily reduced delays but failed to improve compliance. In unseen scenarios, the proposed method consistently reduced delays, while gains in compliance varied with traffic conditions. These results show that integrating violation-sensitive reward shaping into PPO enables policies that minimize both waiting time and violations, offering a practical and robust approach for intelligent traffic signal control in complex urban environments.

D. C. Khrisne, Made Sudarma, I. Giriantari et al. · 0 citations
Open access Sep 2026

Hybrid-RL-RB: A Constraint-Aware Reinforcement Learning and Rule-Based Algorithm for Multi-Intersection Traffic Signal Control

Traffic signal control plays a critical role in mitigating congestion and improving urban mobility, particularly in multi-intersection networks where fixed-time strategies cannot adapt to fluctuating demand. Although reinforcement learning has shown strong potential for adaptive signal optimization, purely learning-based controllers often rely on reward shaping rather than explicit enforcement of traffic engineering constraints, which may lead to unstable phase switching and operational inefficiencies. This study proposes a Hybrid Reinforcement Learning and Rule-based algorithm (Hybrid-RL-RB), a constraint-aware traffic signal control algorithm that combines reinforcement learning with a rule-based supervisory layer for multi-intersection traffic signal control. In the implemented version, the learning component is based on tabular Q-learning with a discretized traffic state representation, while the rule-based layer supervises the final executable signal action. The objective is to improve adaptive signal control while preserving operational feasibility through minimum green time, maximum green time, spillback protection, and phase-safety constraints. The framework was implemented in SUMO through TraCI and evaluated under three scenarios of low, medium, and high traffic demand conditions across multiple network configurations, including a real-network topology (Casablanca-OSM). Experimental results show that Hybrid-RL-RB reduces average queue length by up to 51.47% and waiting time by up to 68.10% compared with Fixed-Time control. Compared with Simple-RL, the proposed method provides modest but consistent queue reductions on the 16 × 16 network, while MaxPressure remains the strongest queue-minimization baseline. In the high-demand Casablanca-OSM scenario, Hybrid-RL-RB reduces queue length by 20.50%, reduces waiting time by 21.41%, and increases throughput by 16.83% compared with Fixed-Time control. These results indicate that explicit rule-based projection can improve the operational feasibility and extensibility of RL-based traffic signal control, although further validation with additional seeds and longer real-network simulations is required.

Mohammed El Kaim Billah, Mohammed-Alamine El Houssaini, Abedelfettah Mabrouk et al. · 0 citations
#reinforcement learning Open access Sep 2026

A framework for benchmarking traffic signal control robustness under incidents: comparative study of reinforcement learning-based methods

Reinforcement learning-based traffic signal control (RL-TSC) has emerged as a promising approach for improving urban mobility. However, its robustness under real-world disruptions such as traffic incidents remains largely underexplored. In this study, we introduce T-REX, an open-source, SUMO-based simulation framework for training and evaluating RL-TSC methods under dynamic, incident scenarios. T-REX models realistic network-level performance considering drivers’ probabilistic rerouting, speed adaptation, and contextual lane-changing, enabling the simulation of congestion propagation under incidents. To assess robustness, we propose a suite of metrics that extend beyond conventional traffic efficiency measures. Through extensive experiments across synthetic and real-world networks, we showcase T-REX for the evaluation of several state-of-the-art RL-TSC methods under multiple real-world deployment paradigms. Our findings show that while independent value-based and decentralized pressure-based methods offer fast convergence and generalization in stable traffic conditions and homogeneous networks, their performance degrades sharply under incident-driven distribution shifts. In contrast, hierarchical coordination methods tend to offer more stable and adaptable performance in large-scale, irregular networks, benefiting from their structured decision-making architecture. However, this comes with the trade-off of slower convergence and higher training complexity. These findings highlight the need for robustness-aware design and evaluation in RL-TSC research. T-REX contributes to this effort by providing an open, standardized and reproducible platform for benchmarking RL methods under dynamic and disruptive traffic scenarios.

Dang Viet Anh Nguyen, Carlos Lima Azevedo, Tomer Toledo et al. · 0 citations
Open access Aug 2026

Edge-Aware Hybrid-Action Reinforcement Learning for Latency-Sensitive Cooperative Bus Signal Priority in Vehicular Edge Networks

Dense short-block road networks require bus signal priority (BSP) decisions to be generated and delivered within short and reliable control windows. This paper presents Edge-HyAR-BSP, a deadline-aware cooperative BSP framework that supports edge-side execution in vehicular edge networks. Roadside edge nodes perform decentralized low-latency inference, while the cloud supports centralized training and model updates. The framework represents each priority decision as a coupled phase-duration action and checks its executability under bus ETA, signal-safety, compensation, and edge-side deadline constraints. Green extension, red truncation, and cross-cycle compensation are integrated to improve bus passage while limiting disturbance to general traffic. The project platform covers 102 signalized intersections in the Rongdong District of Xiong'an New Area. Detailed operational evaluation is conducted on a 15-intersection corridor served by Route 302, while robustness and ablation analyses are performed in simulation using a topology derived from the 102-intersection network. In the selected pre-/post-deployment periods, bus speed increased by 14.64%-30.09%, aggregate bus delay decreased by 38.28%-65.13%, and bus stops decreased by 42.01%-58.62%. Edge deployment reduced mean end-to-end decision latency from 89.7 ms to 24.6 ms. Simulation results show gradual degradation under increased delay, packet loss, and workload. The findings provide operational case-study evidence for the feasibility of edge-aware hybrid-action BSP, while broader multi-route and cross-city validation remains necessary.

Hui Deng, Xiaofei Wang, Zipeng Wang et al. · 0 citations
Preprint Aug 2026

SIGMA: Symmetry-aware, Intelligent, Geometric, Multi-objective Adaptive Control for Robust, Dependable Traffic Management

SIGMA converts natural-language emergency commands into priority vectors for a multi-objective actor-critic controller, avoiding manual reward engineering and offers a reliable, language-guided, multi-objective traffic control system with statistical reliability assurance.

Pratham Payra, B. Jagadish, T. Sen et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.