Reinforcement Learning–Based Operational Optimization of a Hybrid Wind–CSP System with Thermal and Battery Energy Storage
Hybrid renewable power plants that combine multiple generation technologies with energy storage can improve renewable utilization and reliability, but their operation requires solving a high-dimensional, sequential dispatch problem under uncertainty. This paper presents a reinforcement learning (RL) framework for operational dispatch of a hybrid concentrated solar power (CSP)–wind system–Photo Voltaic (PV) solar system equipped with thermal energy storage (TES) and a battery energy storage system (BESS). We reformulate operational dispatch as a Markov decision process (MDP) with continuous control actions and a feasibility-projection layer that maps agent actions to physically admissible power flows. Proximal policy optimization (PPO) is adopted to learn a closed-loop dispatch policy that can react to variability in demand and renewable output. Using an 8064-hour held-out evaluation horizon, PPO with a nominal loss-of-load penalty improves the shaped return by 4.80% relative to a fixed-priority rule-based baseline, but increases loss of power supply probability (LPSP) from 3.60% to 11.17% due to near-elimination of BESS cycling. Increasing the loss-of-load penalty recovers the baseline reliability (LPSP 3.60%) and curtailment (42.25%), but causes PPO to collapse to the rule-based policy. These results highlight the sensitivity of RL dispatch to reward weights and motivate constrained/safe-RL formulations that enforce reliability targets while optimizing storage usage.