Telemetry-Robust Safe Zero-Shot Graph Multi-Agent Reinforcement Learning for Active Voltage Control Across Distribution Feeders
Abstract
Active voltage control in distribution networks depends on a sensing–decision–actuation chain that must remain effective as feeder topology, inverter participation, and telemetry quality change. Most multi-agent reinforcement learning controllers retain feeder-specific observation and action interfaces, and their behavior under imperfect telemetry is rarely tested under whole-graph transfer. This paper proposes GRAS-AVC, which is a zero-shot graph actor–critic framework with a permutation-equivariant shared actor, variable-size twin graph critics, droop-residual actions, and deployment-time AC power-flow risk screening. The 322-bus target feeder contributes no replay, gradient updates, risk fitting, or checkpoint selection. On an independent third-year ten-day test, GRAS-AVC reduces violating bus–time pairs from 7.455% under droop control to 2.723% (63.5% relative reduction). Against a matched edge-conditioned Graph-TD3 backbone, the GRAS actor lowers pooled exposure from 9.239% to 4.027%; after identical screening, GRAS-AVC lowers it from 3.896% to 2.723% while triggering 15.89 percentage points less often. Across information-matched tests spanning reconstructed missing telemetry, graph-correlated errors, gross bad data, a one-step delay, compound stress, and 50 paired announced reconfigurations, the frozen system maintains 63.1–63.6% droop-relative reductions and improves bus–time exposure in every seed. A selector-model audit records no false acceptance across 40 exact-and-bounded-mismatch condition–seed evaluations. GRAS-AVC therefore couples topology-aware policy inference with auditable physical screening for scalable sensing-to-control operation in DER-rich distribution networks.