Skip to content
Open access

A Bilevel Reinforcement Learning Framework for Coordinated EV Charging and Dynamic Pricing

Jul 2026 · Energies · 0 citations · 40 references

TL;DR

Results validate the effectiveness of bilevel RL for complex energy optimization problems, highlighting its potential as a scalable control paradigm for smart management systems.

Abstract

This paper presents a bilevel Reinforcement Learning (RL) framework for optimizing Electric Vehicle (EV) charging through price-mediated coordination between grid operators and charging stations. Unlike prior work relying on direct control or manual subgoal engineering, the proposed approach uses dynamic pricing as an implicit coordination signal to address a complex multi-objective optimization problem involving grid stability, user satisfaction, and economic efficiency. To manage this complexity, the problem is decomposed into two levels comprised of an upper-level Distribution System Operator (DSO) that determines dynamic pricing strategies, and multiple lower-level Load Aggregators (LAs) responsible for EV charging decisions at individual stations in response to these prices. This bilevel structure captures the leader–follower interaction between DSOs and LAs, with each level operating at different temporal scales. Deep Deterministic Policy Gradient (DDPG) agents are deployed at both levels, enabling adaptive decision-making under operational constraints. Extensive simulations compare the framework against multiple Rule-Based Control (RBC) baselines. Results demonstrate that the DDPG-based DSO achieves a 42.4% higher mean reward and 19.1% higher profit compared to the best-performing RBC baseline, while preserving grid stability and user satisfaction. These results validate the effectiveness of bilevel RL for complex energy optimization problems, highlighting its potential as a scalable control paradigm for smart management systems.

Read PDF

Similar papers

Open access Aug 2026

Coordinated Optimization of Orderly Charging and Grid Interaction at Electric Vehicle Charging Stations Based on Multi-Agent Reinforcement Learning

A collaborative optimization framework based on multi-agent reinforcement learning is proposed for orderly charging at electric vehicle charging stations and coordinated interaction with the power grid, providing a technical reference for intelligent charging coordination under grid interaction and electromagnetic comp...

Y.-X. Wang · 0 citations
Jul 2026

A Hierarchical Stochastic Model Predictive Control Framework for Integrated Request-aware Charge Scheduling and Service Allocation

Evaluation of this approach using a state-of-the-art commercial solver with stochastic EV rental requests under different confidence levels and time-varying electricity prices demonstrates significant benefits of the integrated mechanism design, including a reduction in charging costs and battery capacity degradation c...

Mainak Dan, A. Easwaran · 0 citations
Open access Aug 2026

Multi-Objective Reinforcement Learning for Smart Planning of Electric Vehicle Charging Stations

A hybrid optimization framework that combines greedy initialization with reinforcement learning to efficiently explore the charging station deployment problem is proposed and demonstrates stable performance across three evaluated deployment scenarios, indicating its potential applicability to increasingly complex charg...

A. Bousia · 0 citations
Open access Aug 2026

Constructing the Optimal Bidding Path for VPPS Participating in the Spot Market Using the DDPG Reinforcement Learning Algorithm

This study proposes an optimal bidding path construction framework based on the Deep Deterministic Policy Gradient (DDPG) reinforcement learning algorithm for VPP participation in electricity spot markets and provides valuable insights into communicationenabled power systems, distributed electromagnetic information net...

Peng-Tao Hu, N. Shen, H. Guo et al. · 0 citations
Open access Jul 2026

Adaptive primal–dual Q-learning for electric vehicle route optimization on real-world charging networks

A novel Dual Q–Adaptive Weighting model that balances reward and cost through a primal–dual learning mechanism is proposed and achieves the highest route accuracy of 78.66%, outperforming Double Q-Learning, Q-Learning, and traditional A* and Dijkstra algorithms.

Sarvesh Kumar, Rayappa David Amar Raj, Archana Pallakonda et al. · 0 citations
Open access Sep 2026

Coordinated Scheduling of Distribution Network and Transportation System for EVs with Aggregated Flexibility and Endogenous Dynamic Pricing

With the large-scale integration of Electric Vehicles (EVs) into distribution systems, the spatiotemporal uncertainty of charging loads and the interplay between user charging behavior and network operational constraints present new challenges to the safe and economical operation of the power system. To address the ins...

Si-Zu Hou, Yao Sang, Xuan Zhao et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.