Fusion of adaptive decoupling control and deep reinforcement learning for intelligent waste heat recovery
Abstract
Efficient recovery of low-grade industrial waste heat is important for improving energy utilization; however, Roots-expander-based waste heat recovery systems remain difficult to control because of strong multi-parameter coupling, nonlinear operating characteristics, model-switching transients, and stochastic gas-source disturbances. To address these challenges, this study proposes an intelligent hybrid control strategy, termed NMMC-DRL, for a Roots-expander-based waste heat recovery system. The main novelty of the proposed method is that deep reinforcement learning (DRL) is not used as a standalone controller, but as an upper-level residual compensation mechanism embedded in a nonlinear multi-model adaptive decoupling control (NMM-ADC) framework. In this architecture, the lower-level NMM-ADC controller handles nominal nonlinear coupling dynamics and preserves the stabilizing role of the model-based control structure, whereas the upper-level DRL agent, trained using the deep deterministic policy gradient algorithm, learns real-time compensatory actions for unmodeled dynamics, coupling residuals, model-switching effects, and unknown gas-source disturbances. Furthermore, an extended state representation and a compensation-oriented reward function are designed to incorporate tracking errors, coupling residual information, and switching-related effects into the learning process, enabling smoother and more adaptive compensation. The proposed NMMC-DRL strategy is evaluated through comprehensive simulations and physical experiments on a Roots-expander-based waste heat recovery platform. Under random gas-source fluctuations, the controller limits speed variations to within ±2.5%, reduces settling time by approximately 40%, and decreases overshoot by approximately 60%. These results demonstrate the effectiveness of the proposed hybrid control strategy in improving disturbance rejection, tracking accuracy, and operational robustness.