Open access
Aug 2026
Research on Optimal Power Grid Scheduling Based on Transfer Reinforcement Learning
M3-PPO introduces two key innovations to overcome MAML’s training instability: a Mamba-based context encoder for richer task representation in the inner loop, and a global-local momentum update mechanism for smoother meta-parameter optimization in the outer loop.
Q.-H. Dai, X. Hu, J.-L. Li et al.
· Advanced Electromagnetics · 0 citations