Skip to content
Open access

Autonomous Volt/Var Control in Active Distribution Networks via LLM-Driven Dynamic Reward Shaping

Aug 2026 · Electronics · 0 citations · 29 references

TL;DR

A hierarchical autonomous control framework featuring large language model-driven dynamic reward shaping (LLM-Driven DRS) is introduced to balance security and efficiency and achieves Pareto superiority over conventional static weight strategies.

Abstract

To alleviate the severe voltage security and operational efficiency challenges brought about by the increasing penetration of distributed energy resources in active distribution networks, Volt/Var control (VVC) has become a key mechanism to stabilize node voltage and minimize power loss by coordinating reactive power injection. While multi-agent reinforcement learning (MARL) offers a promising decentralized control approach, its static reward functions are prone to creating harsh trade-offs between voltage constraint enforcement and cost-efficiency. In this paper, a hierarchical autonomous control framework featuring large language model-driven dynamic reward shaping (LLM-Driven DRS) is introduced to balance security and efficiency. The dynamic priority shifting (DPS) mechanism lies at the center of the framework and dynamically varies the reward weights through the detection of real-time grid bottlenecks. Under the LLM-Driven DRS framework, this mechanism successfully achieves a fluid transition between a Constraint-Dominant Phase for voltage stabilization and an Objective-Refinement Phase for economic optimization. Validation on a modified IEEE 33-bus system demonstrates that the proposed framework achieves Pareto superiority over conventional static weight strategies. Crucially, compared with the 1.250% static baseline, the absolute voltage violation rate is suppressed to 0.014%, mitigating long-tail risks of hardware degradation and inverter tripping, while active power losses are reduced by up to 33.76%. A robust safety margin is further confirmed by spatiotemporal analysis, which reveals an average minimum voltage margin increase of over 0.011 p.u. under severe stress conditions.

Read PDF

Similar papers

Preprint Aug 2026

Learning to Run Power Networks: Effective AlphaZero-inspired Topological Control

As the integration of volatile renewable energy sources increases the strain on modern power grids, the use of Reinforcement Learning (RL) for autonomous topological reconfiguration has emerged as a promising research field to keep strained grids stable and operational. Compared to traditional redispatching measures, topological actions offer a cheaper and more cost-effective way to manage grid congestion. However, their implementation is hindered by a vast combinatorial action space and strict operational constraints. This paper investigates the effectiveness of model-based AlphaZero-inspired approaches that utilize Monte Carlo Tree Search (MCTS) for proactive grid management. We systematically evaluate how reward functions, observation density, and search guidance influence an agent's survivability. Our results demonstrate that the optimized AlphaZero approach achieves a peak survivability of 98.43%, significantly outperforming the proximal policy optimization (PPO) variant. We find that conducting the MCTS without guidance from a prior learned policy or value function can enhance training efficiency, and that a straightforward binary survival reward provides more effective search guidance than complex, multi-objective functions. Our findings demonstrate that while AlphaZero is a powerful framework for topological control, pure reinforcement learning is not sufficient; rather, an effective and reliable system requires a'minimalist'integration of domain-specific heuristics, binary rewards, and a restricted observation space of line loads.

Lukas Zetto, B. Schäfer, Qiong Huang · 1 citation
Conference Open access Aug 2026

Implementing Dynamic Virtual Power Plants with Model Predictive Control

Power systems are integrating more distributed energy resources (DERs) to meet decarbonization targets. Yet inverter-based generation reduces system inertia and increases the need for fast-acting dynamic ancillary services. Virtual power plants (VPPs) aggregate heterogeneous DERs to provide such services. However, existing approaches do not directly combine prescribed dynamic responses for fast frequency and voltage regulation with explicit enforcement of device-and distribution-network constraints in co-located VPPs. This paper proposes a model predictive control (MPC) framework for dynamic VPPs that tracks grid code-specified behaviour encoded by a desired transfer function while enforcing device-and feeder-level constraints. The framework implements a non-uniform prediction horizon that preserves fine near-term resolution for fast disturbance response while extending look-ahead without uniformly increasing computational burden. Case studies on a modified IEEE 33-bus feeder demonstrate close tracking of frequency and voltage regulation targets with practical real-time feasibility under suitable disturbances, and graceful degradation when requests exceed VPP capacity. The modular design accommodates diverse grid codes and DER portfolios, positioning the framework as a practical tool for evolving ancillary service markets.

Niko Andrianos, Babak Ghaffarzadeh, Dominic Liao-McPherson · 0 citations
Open access Jul 2026

Expert-guided optimization for load transfer in distribution networks assisted by virtual power plants

The rapid expansion of distribution networks and the increasing complexity of their topological structures pose significant challenges to fast and reliable post-fault service restoration. Meanwhile, driven by carbon neutrality goals, the large-scale integration of distributed energy resources (DERs) enhances operational flexibility but also introduces pronounced intermittency and uncertainty, further complicating post-fault load transfer decision-making. To address these challenges, this paper proposes an expert-guided and virtual power plant (VPP)-assisted load transfer optimization framework based on hierarchical graph reinforcement learning. A topology-aware graph neural network (GNN)–based state representation is developed, in which buses are modeled as nodes and switches as controllable edges, enabling explicit modeling of network connectivity and electrical coupling. On this basis, a hierarchical decision-making architecture is constructed: the upper-level agent, guided by expert knowledge, dynamically selects the restoration task type to coordinate the timing of network reconfiguration and VPP-assisted DER regulation; driven by this high-level directive, two specialized lower-level agents respectively execute the specific switch operations and stepwise DER power adjustments, ensuring power balance and voltage security. Simulation results on a practical distribution network demonstrate that, under high DER penetration, the proposed method achieves faster service restoration, higher load recovery ratios, and significantly fewer voltage violation events than conventional reinforcement learning approaches, exhibiting improved operational safety and scheduling stability.

Lu Chen, Jinhu Fang, Xiaona Lv et al. · 0 citations
Open access Sep 2026

Telemetry-Robust Safe Zero-Shot Graph Multi-Agent Reinforcement Learning for Active Voltage Control Across Distribution Feeders

Active voltage control in distribution networks depends on a sensing–decision–actuation chain that must remain effective as feeder topology, inverter participation, and telemetry quality change. Most multi-agent reinforcement learning controllers retain feeder-specific observation and action interfaces, and their behavior under imperfect telemetry is rarely tested under whole-graph transfer. This paper proposes GRAS-AVC, which is a zero-shot graph actor–critic framework with a permutation-equivariant shared actor, variable-size twin graph critics, droop-residual actions, and deployment-time AC power-flow risk screening. The 322-bus target feeder contributes no replay, gradient updates, risk fitting, or checkpoint selection. On an independent third-year ten-day test, GRAS-AVC reduces violating bus–time pairs from 7.455% under droop control to 2.723% (63.5% relative reduction). Against a matched edge-conditioned Graph-TD3 backbone, the GRAS actor lowers pooled exposure from 9.239% to 4.027%; after identical screening, GRAS-AVC lowers it from 3.896% to 2.723% while triggering 15.89 percentage points less often. Across information-matched tests spanning reconstructed missing telemetry, graph-correlated errors, gross bad data, a one-step delay, compound stress, and 50 paired announced reconfigurations, the frozen system maintains 63.1–63.6% droop-relative reductions and improves bus–time exposure in every seed. A selector-model audit records no false acceptance across 40 exact-and-bounded-mismatch condition–seed evaluations. GRAS-AVC therefore couples topology-aware policy inference with auditable physical screening for scalable sensing-to-control operation in DER-rich distribution networks.

Bo-Yin Jin, Si-Qi Sun, Yun Zhang et al. · 0 citations
Open access 2026

Control-Barrier-Function Shielded Conservative Distributional Reinforcement Learning for Anti-Misoperation AGC of Thermal Power Units

: High shares of variable renewable generation increasingly require thermal units to provide deeper, faster, and more frequent automatic generation control (AGC) while operating close to low-load, steam-pressure, ramp-rate, and actuator limits. Under these conditions, an area-level regulation request that is valid for grid balancing can become unsafe at the plant interface because of nonlinear boiler–turbine dynamics, delayed measurements, sensor bias, and compound disturbances. This study therefore formulates AGC from the plant side: an area-level dispatcher allocates regulation among thermal generation, hydro generation, battery storage, and flexible load, while the proposed controller modifies only the command assigned to a 600 MW reheat thermal unit before it enters the plant distributed control system. A control-barrier-function (CBF) shielded conservative distributional reinforcement learning method is developed to prevent excessive ramping, steam-pressure excursions, valve over-actuation, low-load instability, and errors caused by delayed or biased measurements. Every transition used for training, validation, and testing is generated numerically by the declared behavior policy and benchmark generator; no measured AGC record or plant-historian data enter the reported results. Asynchronous signals are causally aligned on a 4 s decision grid, and a dedicated distributional safety-cost critic estimates the 0.95 conditional value-at-risk (CVaR) of discounted constraint cost. The nonlinear simulator is additionally cross-checked against the published plant-derived Bell–Åström model at seven operating points. Normalized power-transient NRMSE is 0.021–0.079 with correlation 𝑟 ≥ 0.995 , while pressure-transient NRMSE is 0.084–0.254 with 𝑟 = 0.733 –0.970. Ten independent training seeds and 3500 held-out numerical episodes show that the controller reduces cumulative absolute ACE from 218 to 165 MW min and improves the frequency nadir from −0.073 to −0.041 Hz relative to PI AGC. No hard violation occurs in 525,000 feasible test decisions, giving a one-sided 95% binomial upper bound of 0.00057% on the per-decision violation probability. The safety statement remains conditional on the declared plant model, uncertainty set, robust control-invariant subset, and bounded prediction errors. End-to-end CPU execution requires 6.8 ms on average and 9.6 ms at the 95th percentile.

Mu-Jie Zhang, Ya-Jun Wu, Dong-Sheng Li et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.