Deep Reinforcement Learning (DRL) holds significant promise for achieving human-like Autonomous Vehicle (AV) capabilities, but suffers from low sample efficiency and challenges in reward design. Model-Based Reinforcement Learning (MBRL) offers improved sample efficiency and generalizability compared to Model-Free Reinforcement Learning (MFRL) in various multi-agent decision-making scenarios. Nevertheless, MBRL faces critical difficulties in estimating uncertainty during the model learning phase, thereby limiting its scalability and applicability in real-world scenarios. Additionally, most studies on Connected Autonomous Vehicles (CAVs) focus on single-agent decision-making. In contrast, existing multi-agent MBRL solutions lack computationally tractable algorithms with Probably Approximately Correct (PAC) guarantees, a crucial factor for ensuring policy reliability with limited training data. To address these challenges, we propose MA-PMBRL, a novel Multi-Agent Pessimistic Model-Based Reinforcement Learning framework for CAVs, incorporating a max-min optimization approach to enhance robustness and decision-making. To mitigate the inherent subjectivity of uncertainty estimation in MBRL and avoid incurring catastrophic failures in AV, MA-PMBRL employs a pessimistic optimization framework combined with Projected Gradient Descent (PGD) for both model and policy learning. MA-PMBRL also employs general function approximations under partial dataset coverage to enhance learning efficiency and system-level performance. By bounding the suboptimality of the resulting policy under mild theoretical assumptions, we successfully establish PAC guarantees for MA-PMBRL, demonstrating that the proposed framework represents a significant step toward scalable, efficient, and reliable multi-agent decision-making for CAVs.
Ruo-Qi Wen, Rongpeng Li, Xing Xu et al.· IEEE Transactions on Mobile...· 1 citation
This work investigates unsourced random access (URA) over quasi-static fading channels, which is well suited for massive connectivity. In most existing URA schemes for fading channels, message recovery typically requires channel estimation, and these schemes therefore rely on pilot-assisted designs, resulting in limited spectral efficiency and degraded performance. In this context, on–off division multiple access (ODMA) has emerged as a promising URA framework due to its super-sparse structure; however, existing ODMA-based schemes still rely on a single large pattern book and exhibit high computational complexity. To overcome these limitations, this paper proposes a novel URA scheme that integrates signal scrambling with an ODMA-based framework, referred to as SS-ODMA. The key idea of SS-ODMA is to introduce a user-specific polarity scrambling pattern as an auxiliary signature, which decouples the original large pattern book into two compact components and mitigates the exponential growth in detection complexity. Building on this structure, we develop a low-complexity three-stage iterative detection algorithm that enables activity detection, joint channel estimation and scrambling pattern detection, and multi-user data recovery without relying on pilot signals. Simulation results demonstrate that the proposed SS-ODMA scheme outperforms existing URA schemes over a wide range of active users while reducing computational complexity by approximately two orders of magnitude.
Jianxiang Yan, Ying Li, Guanghui Song et al.· IEEE Transactions on Wireles...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.