Skip to content

Optimized adaptive traffic management using double DQN and prioritized experience replay in a multi-agent traffic control setup

Jul 2026 · Cluster Computing · Vol 29 · 0 citations · 49 references

TL;DR

Simulation results in the SUMO environment demonstrate that DDQNTSCA achieves faster convergence, enhanced adaptability, and significant reductions in average travel time, queue length, and cumulative delay compared to existing DRL-based TSC methods.

View source

Similar papers

Open access Jul 2026

Dynamic traffic signal scheduling system based on adaptive quad agent Double Deep Q -network algorithm

Real-time estimation of vehicle queue lengths at signalized intersections remains a significant challenge, particularly when conventional input–output traffic models fail to capture queues extending beyond detector coverage. Although Deep Q-Networks (DQNs) have demonstrated considerable potential for dynamic traffic signal control, existing approaches often suffer from large state spaces, unstable reward signals, and inefficient utilization of high-quality traffic data. To address these limitations, this study proposes an Adaptive Quad-Agent Double Deep Q-Network (AQDDQN) framework for intelligent traffic signal optimization. The proposed AQDDQN framework improves learning stability and Q-value estimation accuracy through a multi-agent reinforcement learning strategy. The model analyzes the relationship between vehicle queue length and reward values to optimize signal control decisions. Historical traffic data are utilized to establish preconditions, time-based prediction errors are computed, and optimal signal phases are selected based on minimum loss across multiple preconditions. The experimental evaluation includes agent-wise behavioral analysis and comparative assessments against Double Deep Q-Network (DDQN), Fixed Point Techniques, and Improved DDQN methods. Performance is evaluated using metrics such as reward values, queue lengths, predicted overflow delays, and queue length–reward relationships. The proposed adaptive framework demonstrates superior performance compared with existing approaches by improving traffic signal control accuracy, reducing prediction errors, and enhancing overall traffic throughput. Simulation results indicate that the AQDDQN model effectively supports dynamic signal phase adaptation, minimizes congestion, and provides more accurate queue length estimations under complex traffic conditions. The findings confirm the effectiveness and robustness of the proposed AQDDQN framework for real-time intelligent traffic management. By improving learning stability and adaptive decision-making capabilities, the model offers a practical solution for optimizing traffic operations at signalized intersections and has strong potential for deployment in future smart transportation systems.

Bharathi Ramesh Kumar, Sachin Salunkhe, S. Shinde et al. · 0 citations
Open access Aug 2026

Deep reinforcement learning-based traffic signal control in multi-intersection environments: a comparative study of DQN variants

Traffic congestion at urban intersections is commonly associated with non-adaptive Fixed Time Signal Control (FTSC), which cannot respond effectively to variations in vehicle types, traffic demand, and intersection policies. Although deep reinforcement learning (DRL) has been increasingly applied to traffic signal control, comprehensive evaluations under multi-intersection environments with heterogeneous vehicles and different turning policies remain limited. This study evaluates four value-based DRL algorithms, namely DQN, DDQN, Dueling DQN, and Dueling DDQN, for optimizing traffic signal control in a simulated four-intersection network. The simulation incorporates heterogeneous vehicle types, priority-weighted vehicles, and two turning policy scenarios, and the revised evaluation also includes Fixed Time Signal Control, Longest Queue First, and Max Pressure as baseline controllers. Results from repeated training and testing evaluations show that the DRL-based controllers generally outperform FTSC and remain competitive against adaptive baselines across waiting time and speed metrics. Case 2, which allows direct left turns, consistently performs better than Case 1; however, this improvement is interpreted as the combined effect of DRL-based control and a less restrictive traffic policy. In offline testing for Case 2, Dueling DQN reduces ambulance waiting time from 167.6 to 42.2 s, corresponding to a 74.83% reduction. Overall, the findings demonstrate the potential of DRL-based traffic signal control in controlled simulation conditions and highlight that algorithm performance is strongly influenced by traffic policy design and environmental complexity.

D. Prastiyanto, A. A. Manaf, Muhammad Ahnaf Maulana et al. · 0 citations
Review Open access Aug 2026

Reinforcement learning for traffic signal control in large-scale transportation networks: a systematic literature review

Effective Traffic Signal Control (TSC) in large-scale transportation networks is essential for enhancing urban mobility, reducing congestion, and improving safety. However, traditional control methods often fail to effectively address the complexity, dynamic conditions, and multimodal demands of modern urban traffic systems. In recent years, Reinforcement Learning (RL) has emerged as a promising solution for achieving adaptive and scalable TSC. This paper presents a systematic and up-to-date review of RL-based methods for large-scale TSC. We analyze representative studies published between 2013 and 2025, presenting a comprehensive analysis of traffic simulation environments, transportation modalities, and advances in methodologies. Key aspects include multi-agent paradigms, state and action representations, reward mechanisms, RL frameworks, as well as advanced representation learning and cooperative strategies for large-scale transportation networks. We also provide a critical discussion on performance evaluation and opportunities for improvement, and conclude by summarizing the current challenges and outlining future research directions. This review aims to inform and guide the development of next-generation RL-based TSC systems that promote sustainable, safe, and efficient urban transportation.

Xiaocai Zhang, Zhe Xiao, Tao Liu et al. · 0 citations
Aug 2026

Multi-Intersection Traffic Signal Control Based on Multi-agent Reinforcement Learning: A Cooperative Approach

Adaptive traffic signal control has been widely investigated for several decades as an effective solution to mitigate urban traffic congestion. Recently, Multi-agent Reinforcement Learning (MARL) has emerged as a promising approach for optimizing traffic signal control in complex urban networks. However, the decision-making process becomes significantly more challenging in dynamic traffic environments involving multiple interconnected intersections. In such settings, effectively coordinating traffic signals using information collected from the Internet of Things remains a major challenge for improving the overall efficiency of road networks. Despite recent advances, existing MARL-based approaches still exhibit several limitations. First, many methods struggle to effectively integrate heterogeneous traffic information from complex traffic environments. Second, they often neglect the correlations among agents and fail to adequately capture the spatial–temporal dependencies between neighboring intersections. To address these challenges, this paper proposes a novel cooperative MARL-based approach for adaptive traffic signal control in multi-intersection networks. The proposed framework employs double deep Q-network architecture for each agent and introduces a cooperation mechanism that enables agents to share aggregated traffic information from neighboring intersections. Furthermore, a pheromone-based state representation is incorporated to model the collective traffic conditions and enhance coordination among agents. Extensive experiments conducted under different traffic scenarios demonstrate that the proposed approach significantly outperforms existing methods in relation to average pheromone intensity, average noise emission, and average waiting time. These results highlight the effectiveness of the proposed cooperative MARL framework in improving traffic efficiency and reducing congestion in multi-intersection urban traffic networks.

T. Haddad · 0 citations
Conference Jul 2026

Optimization of last-mile delivery using dynamic path algorithms based on reinforcement learning

Deep reinforcement learning has emerged as a transformative approach for solving the challenges of urban last-mile delivery, characterized by rapidly changing demand and unpredictable traffic conditions. The study uses a dynamic routing framework to create adaptive routing strategies, centered on a customized deep Q-network. By collecting real-time traffic data, vehicle status and delivery priority, the framework can formulate routing strategies. Using advanced data fusion algorithms, the method integrates heterogeneous sensors and traffic inputs. It also creates a composite reward function to balance operational costs, on-time delivery, and service quality. Simulations in a high-fidelity urban environment show that the proposed system can reduce the total delivery cost by 15 to 20 percent compared to static routing methods, while improving energy utilization and time efficiency. Sensitivity analysis shows that the method works well in variable traffic scenarios and effectively adapts to peak and off-peak delivery windows. These findings suggest that deep reinforcement learning and real-time data integration, along with multiobjective policy design, are effective solutions for optimizing modern urban logistics and last-mile delivery networks.

Qian Zhou · 0 citations