Skip to content
Preprint

SIGMA: Symmetry-aware, Intelligent, Geometric, Multi-objective Adaptive Control for Robust, Dependable Traffic Management

Aug 2026 · 0 citations · 58 references
Computer Science Mathematics

TL;DR

SIGMA converts natural-language emergency commands into priority vectors for a multi-objective actor-critic controller, avoiding manual reward engineering and offers a reliable, language-guided, multi-objective traffic control system with statistical reliability assurance.

Abstract

Traffic signal control is a complex sequential decision-making problem requiring real-time adaptation and trade-offs among throughput, delay fairness, signal stability, and emergency vehicle priority. Existing RL methods often fix objectives, ignore dynamic priority changes, and fail to generalize across geometrically similar intersections.We propose SIGMA (Symmetry-aware, Intelligent, Geometric, Multi-objective Adaptive traffic control), an RL framework enhanced with a large language model (LLM) for adaptive objective tuning and orientation-invariant learning. SIGMA converts natural-language emergency commands into priority vectors for a multi-objective actor-critic controller, avoiding manual reward engineering. Rotational augmentation improves transferability across four-way intersections, while offline-to-online learning ensures stable initialization and gradual adaptation to changing traffic.We define reliability properties covering emergency service levels, graceful degradation under LLM failures, and demand sensitivity, validated via bootstrap statistics. Evaluated in SUMO on four Kolkata-based urban intersections against fixed-time, actuated, and DQN controllers, SIGMA reduces average/emergency waiting times and queue lengths, and boosts throughput. Ablation studies confirm robustness to component failures and geometric rotations. Overall, SIGMA offers a reliable, language-guided, multi-objective traffic control system with statistical reliability assurance.

View source

Similar papers

Open access Sep 2026

Hybrid-RL-RB: A Constraint-Aware Reinforcement Learning and Rule-Based Algorithm for Multi-Intersection Traffic Signal Control

Traffic signal control plays a critical role in mitigating congestion and improving urban mobility, particularly in multi-intersection networks where fixed-time strategies cannot adapt to fluctuating demand. Although reinforcement learning has shown strong potential for adaptive signal optimization, purely learning-based controllers often rely on reward shaping rather than explicit enforcement of traffic engineering constraints, which may lead to unstable phase switching and operational inefficiencies. This study proposes a Hybrid Reinforcement Learning and Rule-based algorithm (Hybrid-RL-RB), a constraint-aware traffic signal control algorithm that combines reinforcement learning with a rule-based supervisory layer for multi-intersection traffic signal control. In the implemented version, the learning component is based on tabular Q-learning with a discretized traffic state representation, while the rule-based layer supervises the final executable signal action. The objective is to improve adaptive signal control while preserving operational feasibility through minimum green time, maximum green time, spillback protection, and phase-safety constraints. The framework was implemented in SUMO through TraCI and evaluated under three scenarios of low, medium, and high traffic demand conditions across multiple network configurations, including a real-network topology (Casablanca-OSM). Experimental results show that Hybrid-RL-RB reduces average queue length by up to 51.47% and waiting time by up to 68.10% compared with Fixed-Time control. Compared with Simple-RL, the proposed method provides modest but consistent queue reductions on the 16 × 16 network, while MaxPressure remains the strongest queue-minimization baseline. In the high-demand Casablanca-OSM scenario, Hybrid-RL-RB reduces queue length by 20.50%, reduces waiting time by 21.41%, and increases throughput by 16.83% compared with Fixed-Time control. These results indicate that explicit rule-based projection can improve the operational feasibility and extensibility of RL-based traffic signal control, although further validation with additional seeds and longer real-network simulations is required.

Mohammed El Kaim Billah, Mohammed-Alamine El Houssaini, Abedelfettah Mabrouk et al. · 0 citations
Open access Aug 2026

Safety-Aware Reinforcement Learning Model for Adaptive Traffic Signal Optimization in Work Zone Environments

The findings show that a single controller trained with surrogate safety indicators as learning objectives can improve operational performance while reducing safety-critical instability in work zones.

Israel Afriyie, Kwadwo Amankwah-Nkyi, Percy Agyei-Essiful et al. · 1 citation
#artificial intelligence Preprint Sep 2026

From Prior-Guided Heuristics to Deployable Agents: Accelerating Demonstration-Driven Reinforcement Learning for Deadline-Constrained Network Control

Timely delivery of delay-sensitive information over dynamic, heterogeneous networks is essential for NextG interactive applications, yet providing strict End-to-End (E2E) peak latency guarantees remains an open challenge. Two obstacles limit the adoption of learning-based network control in this setting: traditional volume-based routing metrics, while highly effective for general traffic management, are not designed to capture traffic urgency; and Deep Reinforcement Learning (DRL) controllers trained from scratch suffer from sample inefficiency, long training times, and early-stage exploration volatility. This paper introduces a deployment-focused network control framework that addresses both obstacles. First, we present Effective Congestion (EC), a deadline-aware metric family that quantifies interface congestion by packet urgency and proactively filters non-viable traffic, coupled with a Uniform Path Grouping (UPG) distribution heuristic promoting robust load-balancing; the resulting policies are embedded into Multi-Agent Deep Reinforcement Learning Effective Congestion ($p^*$) (MADRL EC ($p^*$)), a hybrid architecture combining a distributed scheduler with a centralized RL-based router. Second, we introduce a unified training objective that generalizes existing policy-learning paradigms---behavioral cloning, offline Reinforcement Learning (RL), online RL, and offline-to-online schemes---as special cases, combining a live-reward term, a pre-collected-reward term, and a policy-imitation term. From this objective, we derive the Model-Guided Annealed Reinforcement Learning (MGA-RL) protocol, instantiated on a Deep Deterministic Policy Gradient (DDPG) backbone: a deployment-oriented, demonstration-driven training approach that generalizes conventional Offline-to-Online (O2O) schemes, in which trajectories from a lightweight [...]

Vincenzo Norman Vitale, Mohammad Solki, A. Tulino et al. · 0 citations
Open access Jul 2026

Robust Offline Multi-Agent Reinforcement Learning for Latency-Aware SDN Path Control in 6G-Oriented Network Softwarization

Future sixth-generation (6G)-oriented networks require programmable control that can adapt routing to latency and congestion without unsafe online exploration. This study evaluates offline multi-agent deep deterministic policy gradient (MADDPG) with behavior-adjusted training rewards for latency-aware path control in software-defined networking (SDN). Each traffic pair is modeled as an agent selecting one of three retained candidate paths, while centralized critics learn coordinated decisions from topology-specific Ryu–Mininet transition datasets. Nine policies are compared using ten paired seeds on fat-tree, mesh-grid, and WAN-corridors topologies under a deployed utilization–latency weighting of 0.60/0.40, together with flow-completion, latency, congestion, architectural-comparison, sensitivity, robustness, statistical, and controller-overhead analyses. The utilization-aware path heuristic achieves the strongest overall reward ranking. MADDPG is the strongest learned policy on fat-tree, is not significantly outperformed by any evaluated policy on mesh-grid, and remains statistically tied with completion-matched policies on WAN-corridors. Behavior adjustment is topology-dependent rather than uniformly beneficial. The exported policy requires approximately 52μs per joint decision, whereas complete control-loop timing is dominated by network-statistics polling. These results support offline multi-agent SDN control as a competitive, low-overhead option when interpreted jointly with topology structure, flow completion, and strong heuristic baselines.

A. Kyzyrkanov, Y. Nurakhov, Zhenis Otarbay et al. · 0 citations
Open access Aug 2026

Deep reinforcement learning-based traffic signal control in multi-intersection environments: a comparative study of DQN variants

The findings demonstrate the potential of DRL-based traffic signal control in controlled simulation conditions and highlight that algorithm performance is strongly influenced by traffic policy design and environmental complexity.

D. Prastiyanto, A. A. Manaf, Muhammad Ahnaf Maulana et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.