Sep 2026· Journal of engineering and applied sciences· Vol 73· 0 citations· 19 references
Microgrid Control and Optimization
Abstract
Microgrid energy management systems (EMS) require real-time, economically optimal, and safe dispatch strategies under high renewable uncertainty. While Deep Reinforcement Learning (DRL) offers promising online decision-making capabilities, standard DRL algorithms struggle with the “cold start” problem, slow convergence, and severe Sim-to-Real performance degradation when transferred from simulation to real-world operating conditions. To address these challenges, this paper proposes a novel Curriculum Learning-enhanced Proximal Policy Optimization (CL-PPO) framework for grid-connected microgrid energy management. By structuring a high-fidelity virtual environment into a three-stage progressive curriculum (Basic, Volatile, and Extreme), the agent systematically evolves from learning fundamental arbitrage rules to mastering robust dispatch under severe boundary conditions. Experimental results demonstrate that curriculum-guided knowledge transfer reduces the required convergence steps by approximately 38% compared with Baseline PPO. In typical-scenario evaluations, CL-PPO reduces daily operating costs by 7.1% relative to Baseline PPO and achieves a near-optimal 2.4% economic gap relative to the non-causal Mixed-Integer Linear Programming (MILP) oracle, requiring only 12.5 milliseconds per inference. Furthermore, an offline Sim-to-Real replay evaluation was conducted using approximately 90,000 previously unseen operational records from a real-world campus microgrid. In a replay environment parameterized according to the target microgrid’s energy-storage-system characteristics, CL-PPO demonstrated strong cross-domain robustness against measurement noise and operational disturbances, achieving an average total constraint violation rate of 1.2%. No policy-generated command was applied to the physical microgrid during this evaluation. Ultimately, this research enhances the interpretability of DRL policy evolution and provides a safety-aware and computationally efficient training framework with promising potential for future physical deployment, subject to further hardware-in-the-loop and on-site closed-loop validation. A curriculum learning-enhanced PPO (CL-PPO) framework is proposed for safety-aware and computationally efficient energy management in grid-connected microgrids. A dual-condition transition criterion jointly evaluates reward and safety violation rate to guide curriculum progression. A decoupled reward function separates economic incentives from safety-boundary penalties, improving policy interpretability. CL-PPO reduces the required convergence steps by approximately 38% and narrows the economic gap relative to the MILP oracle to only 2.4%. Offline Sim-to-Real replay evaluation using approximately 90,000 unseen real-world operational records demonstrates strong cross-domain robustness, with an average total constraint violation rate of 1.2%.
The results indicate that software engineering work practices are chosen opportunistically, adapted and configured to provide value under the constrains imposed by the startup context.
Nicolò Paternoster, Carmine Giardino, M. Unterkalmsteiner et al.· Information and Software Tec...· 394 citations· ⚡54
This state-of-practice investigation was performed using a literature review followed by a multiple-case study approach and presents how inconsistency between managerial strategies and execution can lead to failure by means of a behavioral framework.
Carmine Giardino, Xiaofeng Wang, P. Abrahamsson· International Conference on...· 175 citations· ⚡19
This study conducts a case survey study based on the secondary data of the major pivots happened in 49 software startups, and demonstrates that customer need pivot is the most common among all pivot types.
Sohaib Shahid Bajwa, Xiaofeng Wang, Anh Nguyen-Duc et al.· Empirical Software Engineeri...· 127 citations· ⚡15
It is found that roles of MVPs in startups were not fully aware by entrepreneurs, and entrepreneurs should consider a systematic approach to fully explore the value of MVP, as a multiple facet product (MFP).
Anh Nguyen-Duc, P. Abrahamsson· International Conference on...· 93 citations· ⚡9
It is found that what perceived as biggest challenges by software startups do vary across different life cycle stages, even though its significance decreases when the learning focuses of the startups move from problem to solution and their products mature.
Xiaofeng Wang, Henry Edison, Sohaib Shahid Bajwa et al.· International Conference on...· 62 citations· ⚡6
A comprehensive overview of how enhanced sampling methods are reshaping the field, with a particular focus on the data-driven construction of collective variables, is provided.
Kai Zhu, Enrico Trizio, Jintu Zhang et al.· Chemical Reviews· 58 citations
Related blog posts
MIT News · Artificial Intelligence· news.mit.eduOct 2, 2026
Martin Trust Center Managing Director Bill Aulet introduces Dear Dreamer, a free platform for middle and high school students who want to learn about entrepreneurship.
Microsoft Research Blog· microsoft.comSep 30, 2026
Extreme space-weather events can damage power systems on Earth and degrade GPS accuracy and satellite operations. A new machine learning system can predict where damage is likely to occur 30-60 minutes before a storm arrives. The post Forecasting space weather risks on power grids appeared first on Microsoft Research.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.