Skip to content
Open access

Hybrid-RL-RB: A Constraint-Aware Reinforcement Learning and Rule-Based Algorithm for Multi-Intersection Traffic Signal Control

Sep 2026 · Future Transportation · 0 citations · 23 references

Abstract

Traffic signal control plays a critical role in mitigating congestion and improving urban mobility, particularly in multi-intersection networks where fixed-time strategies cannot adapt to fluctuating demand. Although reinforcement learning has shown strong potential for adaptive signal optimization, purely learning-based controllers often rely on reward shaping rather than explicit enforcement of traffic engineering constraints, which may lead to unstable phase switching and operational inefficiencies. This study proposes a Hybrid Reinforcement Learning and Rule-based algorithm (Hybrid-RL-RB), a constraint-aware traffic signal control algorithm that combines reinforcement learning with a rule-based supervisory layer for multi-intersection traffic signal control. In the implemented version, the learning component is based on tabular Q-learning with a discretized traffic state representation, while the rule-based layer supervises the final executable signal action. The objective is to improve adaptive signal control while preserving operational feasibility through minimum green time, maximum green time, spillback protection, and phase-safety constraints. The framework was implemented in SUMO through TraCI and evaluated under three scenarios of low, medium, and high traffic demand conditions across multiple network configurations, including a real-network topology (Casablanca-OSM). Experimental results show that Hybrid-RL-RB reduces average queue length by up to 51.47% and waiting time by up to 68.10% compared with Fixed-Time control. Compared with Simple-RL, the proposed method provides modest but consistent queue reductions on the 16 × 16 network, while MaxPressure remains the strongest queue-minimization baseline. In the high-demand Casablanca-OSM scenario, Hybrid-RL-RB reduces queue length by 20.50%, reduces waiting time by 21.41%, and increases throughput by 16.83% compared with Fixed-Time control. These results indicate that explicit rule-based projection can improve the operational feasibility and extensibility of RL-based traffic signal control, although further validation with additional seeds and longer real-network simulations is required.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.