2026· Computer Modeling in Engineering & Sciences· Vol 148, pp. 1-10· 0 citations· 25 references
TL;DR
A novel reinforcement learning (RL)-based controller, designed specifically for solar thermal systems, designed to learn directly from operational data, enabling it to adapt its control policy in real time to mitigate external disturbances.
Abstract
: The effective control of parabolic trough collectors (PTCs) remains a significant challenge due to the inherent non-linearities of the system and the continuous impact of environmental disturbances. Although PTCs are a key technology for industrial process heat and large-scale electricity generation, classical control strategies often struggle to maintain optimal performance under fluctuating conditions. To address these limitations, this paper presents a novel reinforcement learning (RL)-based controller, designed specifically for solar thermal systems. The proposed RL agent is designed to learn directly from operational data, enabling it to adapt its control policy in real time to mitigate external disturbances. Experimental results demonstrate that the RL controller achieves a fast and well-damped closed-loop response, significantly outperforming traditional control benchmarks. Specifically, the RL controller is compared with a Proportional-Integral controller combined with a feedforward controller, and with a Model-based Predictive Controller. In all simulation-based comparisons, the RL controller outperforms the aforementioned controllers in terms of setpoint tracking or disturbance rejection. These results highlight the potential of machine learning to improve the operational reliability and efficiency of complex renewable energy systems.
Advanced control of heating and cooling systems can substantially reduce energy costs and pollution. However, real-world adoption of popular algorithms among researchers, such as model predictive control (MPC) and reinforcement learning (RL), remains limited due in part to their high deployment and commissioning costs. Here, we develop two nearly commissioning-free controllers tailored to objectives that depend linearly on the controlled thermal load, such as energy costs and pollution. The controllers require at most two thermal parameters. In representative heating simulations, controller performance is robust to large parameter specification errors, suggesting potential for deployment with no tuning. The controllers maintain good occupant comfort while achieving 43 to 98% (depending on the electricity pricing and controller variant) of the performance improvement achieved by an omniscient policy with perfect model information and forecasts. These results suggest that simple, structure-exploiting controllers may capture most of the attainable value of advanced control while avoiding the data, modeling, tuning, and computational burdens that can arise with conventional MPC or RL.
W. G. Dierking, Arash Khabbazi, L. D. Premer et al.· 0 citations
Precise temperature regulation for the evaporation process in display manufacturing remains a significant challenge due to the nonlinearity and complexity of dynamics, safety constraints, and slow thermal responses. The conventional proportional-integral-derivative (PID) control approach has been widely used due to its simplicity; however, it often fails to maintain stability and accuracy under such nonlinear and time-varying conditions. To overcome these limitations, this paper proposes a reinforcement learning–based temperature control framework tailored for the evaporation process. A data-driven heating simulator is first developed to provide a safe and efficient virtual environment for reinforcement learning (RL) training, eliminating the need for costly and risky high-temperature experiments. A hybrid-mode RL architecture is then introduced, consisting of two agents: one for directly controlling power in the early phase and another for RL-based PID gain adjustment in the later phase. Experimental results from both the simulator and the actual heating system demonstrate that the proposed framework achieves stable temperature control. Moreover, our proposed approach significantly outperforms conventional PID control by reducing delay time by 6.90%, maximum overshoot by 94.82%, and settling time by 36.47%. For nonlinear high-temperature processes, the proposed method successfully closes the gap between RL-based control and conventional PID.
B. Park, Narim Jeong, Hyukjun Yang et al.· IEEE Access· 0 citations
Abstract. The need to achieve sustainable production has become an urgent necessity in the conditions of stricter environmental requirements and the rise in the cost of energy worldwide. The classical proportionalintegralderivative controllers and linear Model Predictive Controllers are conventional model-based control strategies that by nature rely on precise process models and fixed optimization horizons and hence are not well suited to the dynamic, non-linear, and complex nature of the modern production environment. The current paper suggests a new adaptive control system based on Reinforcement Learning (RL), where the overall production system is optimized and controlled to achieve Sustainable Production Systems, with the multi-objective rewarding function, which aims to minimize the Overall Equipment Effectiveness (OEE), defect rate, specific energy consumption, and carbon dioxide emissions, and a physics-informed Digital Twin safety filter that stops unsafe policy execution in training and deployment. The proposed framework is evaluated on a multi-machine flexible manufacturing cell benchmark, which includes CNC milling, turning, and robotic assembly, and yields an OEE of 93.6, a defect rate of 0.9, a decrease in specific energy consumption of 27.8, and a decrease in CO 2 emissions of 23.4 compared to PID baseline controllers. Experiments with ablation prove all the above-mentioned in the necessity of each of the architectural constituents, and the Pareto frontier analysis proves that the proposed SAC agent is the best in the OEE-versus-energy trade-off space compared to all of the competing methods.
Apoorva Verma· Materials Research Proceedin...· 0 citations
Traditional control strategies, such as droop-frequency and PI controllers, often show limited adaptability when the system operates under highly variable renewable generation and load conditions. In order to overcome this limitation, this study proposes the design and simulation of a clustered microgrid supported by reinforcement learning (RL)-based virtual inertia control. Three continuous-control RL algorithms were evaluated: Deep Deterministic Policy Gradient (DDPG), Twin-Delayed Deep Deterministic Policy Gradient (TD3), and Soft Actor-Critic (SAC). The SAC agent provided the most robust training performance, reaching stable convergence after approximately 400 episodes and a final reward close to 86 units after 600 episodes. DDPG presented the second-best behavior, whereas TD3 achieved the lowest final reward, approximately 43 units. The proposed Agent SAC-Reinforcement Learning control was tested on a two-area microgrid cluster, which demonstrated greater frequency stability. Results indicate a frequency nadir of 59.82 Hz and ROCOF of 0.1484 Hz/s, with 6.78% nadir deviation improvement and 37.23% ROCOF reduction compared to PID-based strategies.
Adrián Criollo, D. Benavides, Paul Arévalo-Cordero et al.· Applied Sciences· 0 citations
Magnetic levitation (Maglev) systems are a testbed for advanced control strategies because of their nonlinearity, open-loop instability, and high dynamics. Accurate position control of the levitated object is a difficult task, especially under parameter uncertainties and disturbances. Traditional control approaches, like proportional-integral-derivative (PID) control, are based on linearization of the system and hence suffer from poor performance in the presence of nonlinearities. Sliding mode control (SMC) is a robust nonlinear control method that provides good disturbance rejection and parameter uncertainty tolerance. However, its performance is sensitive to the choice of controller gains, which is often a time-consuming and suboptimal manual task. To overcome this challenge, previous research has used metaheuristic algorithms, like genetic algorithms (GA), for gain tuning. These approaches enhance tracking accuracy and transient response, but are offline and non-adaptive to dynamic system changes. In this work, a Deep Deterministic Policy Gradient (DDPG) reinforcement learning approach is developed for online gain tuning of SMC in a non-linear Maglev system. The DDPG agent engages in continuous interaction with the environment to learn an optimal policy for gain adaptation in response to system dynamics. Asymptotic stability of the closed-loop system is analyzed using Lyapunov’s stability theorem for fixed gains, and the adaptive gain case is discussed. The proposed DDPG-tuned SMC is shown to have better tracking performance, shorter settling time, and better disturbance rejection and parameter uncertainty tolerance than the GA-tuned SMC through simulation studies.
Syed Tariq Naqshbandi, Huzaifah Abbas, D. Zaman· International Conference on...· 0 citations