Comparative Evaluation Of Twin Delayed Deep Deterministic Policy Gradient Reinforcement Learning For Integrated Electrical-Hydrogen Energy Systems
Abstract
This paper presents a simulation-based comparative evaluation of conventional optimization and twin delayed deep deterministic policy gradient (TD3)-based reinforcement learning methods for real-time energy management in an integrated electrical-hydrogen energy system (IEHES). With the increasing penetration of photovoltaic (PV) and wind generation, coordinated dispatch of batteries, electrolyzers, hydrogen storage, fuel cells, and grid exchange becomes challenging due to renewable uncertainty and multi-energy coupling. To address this issue, six controllers, including a rule-based energy management system (EMS), day-ahead optimization (DA-Opt), model predictive control (MPC), TD3, Safe TD3, and hierarchical day-ahead/real-time TD3 (H-TD3), are evaluated under the same 24-hour simulation profiles, device constraints, and operating objective. The benchmark considers operating cost, hydrogen shortage, constraint violation, renewable curtailment, online computation time, and forecast-error robustness. The results show that H-TD3 achieves the lowest operating cost of ${1 6. 1 5} \pm {1 7. 1 4}$ and the highest reward of ${- 1 6. 1 5} \pm {1 7. 1 4}$ among the tested controllers. Safe TD3 and H-TD3 eliminate observed constraint violations, while MPC shows the smallest relative cost increase under forecast errors. These results indicate that TD3-based controllers can complement conventional optimization by providing fast online corrective dispatch, although robustness and economic optimality should be evaluated separately.