Skip to content
Conference

Reinforcement Learning–Based Operational Optimization of a Hybrid Wind–CSP System with Thermal and Battery Energy Storage

Jul 2026 · International Conference on Control, Decision and Information Technologies · pp. 951-956 · 0 citations · 16 references

Abstract

Hybrid renewable power plants that combine multiple generation technologies with energy storage can improve renewable utilization and reliability, but their operation requires solving a high-dimensional, sequential dispatch problem under uncertainty. This paper presents a reinforcement learning (RL) framework for operational dispatch of a hybrid concentrated solar power (CSP)–wind system–Photo Voltaic (PV) solar system equipped with thermal energy storage (TES) and a battery energy storage system (BESS). We reformulate operational dispatch as a Markov decision process (MDP) with continuous control actions and a feasibility-projection layer that maps agent actions to physically admissible power flows. Proximal policy optimization (PPO) is adopted to learn a closed-loop dispatch policy that can react to variability in demand and renewable output. Using an 8064-hour held-out evaluation horizon, PPO with a nominal loss-of-load penalty improves the shaped return by 4.80% relative to a fixed-priority rule-based baseline, but increases loss of power supply probability (LPSP) from 3.60% to 11.17% due to near-elimination of BESS cycling. Increasing the loss-of-load penalty recovers the baseline reliability (LPSP 3.60%) and curtailment (42.25%), but causes PPO to collapse to the rule-based policy. These results highlight the sensitivity of RL dispatch to reward weights and motivate constrained/safe-RL formulations that enforce reliability targets while optimizing storage usage.

View source