Skip to content
Conference

A Hierarchical Reinforcement Learning Framework with Spatial-Temporal Graph Attention for Autonomous Driving Decision-Making and Control

Jul 2026 · International Conference on Control, Decision and Information Technologies · pp. 37-42 · 0 citations · 21 references

Abstract

This paper presents a hierarchical framework that integrates spatial-temporal graph attention network (ST-GAT) with reinforcement learning for decision and control of autonomous driving. Inspired by the principles of human cognition, the framework decomposes the driving task into two complementary levels: a high-level trajectory planning module that utilizes the soft actor-critic (SAC) algorithm within the Frenet coordinate system, and a low-level tracking control module based on the worst-case soft actor-critic (WCSAC) strategy. This hierarchical decomposition improves policy stability and sample efficiency by decoupling strategic trajectory planning from reactive control execution. Unlike previous methods, the proposed ST-GAT module enables explicit scene understanding by modeling surrounding vehicles and their interactions as a spatial-temporal graph structure. Through attention-based aggregation, the system dynamically captures road geometry and agent behaviors directly from online sensor observations, enabling mapless situational reasoning. Experimental results show that the proposed framework achieves success rates of 92.33% and 97.00% in the roundabout and five-way intersection scenarios in the CARLA simulator.

View source

Similar papers

Preprint Sep 2026

Graph-Based Safe Reinforcement Learning for Multi-Agent Systems with Time-Varying Topology

A graph-based safe multi-agent reinforcement learning (MARL) framework for cooperative navigation with time-varying topology is presented, integrating a attention-based actor and a Graph Attention Network (GAT) centralized critic, enabling scale-insensitive policy learning under time-varying communication topologies.

Sizhe Xiao, Li-Jing Dong, Rui-Ting Bai et al. · 0 citations
Open access Aug 2026

Spatio-Temporal Attention-Based Improved MADDPG Algorithm for Multi-UAV Formation Path Planning

A novel joint optimization framework, named spatio-temporal attention-based multi-agent deep deterministic policy gradient (STA-MADDPG), which integrates advanced spatial-temporal feature extraction with heuristic gradient guidance and aims to balance computational complexity and adaptive behavior.

Dong Zhao, Huaizhi Dong, Wen-Jing Ren · 0 citations
Open access Sep 2026

TMG-MADRL: a transformer-based meta-graph multi-agent deep reinforcement learning framework for robot path planning in dynamic environments

A novel Transformer-guided Meta-learning and Graph-enhanced Multi-Agent Deep Reinforcement Learning (TMG-MADRL) framework for intelligent robot navigation to achieve robust path optimization, adaptive decision-making, and proactive collision avoidance in dynamic scenarios.

Shu-Lin Song, Lan Wu · 0 citations
Jul 2026

Attention and Intrinsic Curiosity-Enhanced Deep Reinforcement Learning for Path Planning in Dynamic Environments

An end-to-end path planning framework built upon the Proximal Policy Optimization algorithm that achieves significant improvements in key metrics such as path success rate, travel time, and path efficiency compared to baseline methods like A*+DWA and standard DRL.

Shiquan Shen, Jiahao Liu, Zheng Chen et al. · 0 citations
Jul 2026

Reward-Driven LLM Agent Workflows: Synthesizing POMDP Routing and Self-Correction for Autonomous Decision-Making

This research bridges theoretical foundations of reinforcement learning and graph-based memory with autonomous agent workflows, and offers a practical, scalable reference framework for developing artificial intelligence technologies in complex, multi-step autonomous systems.

Amez Amanj Ali, Kuo-Kun Tseng · 0 citations
Open access 2026

Extensive Exploration in Highway Overtaking Scenarios Using Hierarchical Reinforcement Learning

A hierarchical reinforcement learning framework for autonomous highway driving that decomposes delayed-reward highway overtaking decision making into interpretable subtasks and achieves more reliable trap-escape performance than other hierarchical structures, including h-DQN and HIRO.

Zhihao Zhang, Ekim Yurtsever, K. Redmill · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.