Sep 2026· European Conference on Electrical Engineering and Computer Science· Vol 14327, pp. 143271J - 143271J-16· 0 citations· 13 references
Engineering
TL;DR
This paper proposes a lightweight, simplified Q-learning energy management strategy for extended-range electric vehicles (REEVs), successfully implemented on an STM32 microcontroller, fully satisfying strict on-board embedded system constraints and providing a highly feasible solution for intelligent REEV energy management.
Abstract
To address the poor adaptability of traditional rule-based control, the operational instability of basic Q-learning algorithms, and the critical difficulty of deploying complex reinforcement learning models on resource-constrained on-board embedded platforms, this paper proposes a lightweight, simplified Q-learning energy management strategy for extended-range electric vehicles (REEVs), successfully implemented on an STM32 microcontroller. The algorithm achieves significant computational reduction by simplifying the traditional 5×5 state-action space into a highly condensed 2×2 grid. Furthermore, a power cooling mechanism is introduced, a multi-dimensional reward function is reconstructed to balance competing vehicle demands, and an ε-decay exploration strategy is designed. Software-in-the-loop (SIL) simulation verification demonstrates that the proposed strategy tightly controls the state-of-charge (SOC) standard deviation within 0.09. Additionally, high-frequency power fluctuations and range extender start-stop times are drastically reduced, and overall energy efficiency is improved by 50.3% compared with traditional strategies. The optimized algorithm occupies only 72.3% of RAM and 68.7% of Flash memory, fully satisfying strict on-board embedded system constraints and providing a highly feasible solution for intelligent REEV energy management.
Simulation results demonstrate that, compared to the traditional hierarchical optimization framework, the proposed strategy achieves significant improvements in terms of mean absolute jerk, root-mean-square (RMS) value of acceleration, power demand, battery SOH degradation, battery temperature violation, and comprehens...
Chengrui Zhang, Fei Ju, Sichen Gao et al.· Proceedings of the Instituti...· 0 citations
Experimental results demonstrate that the proposed DRL-based framework provides an effective and scalable solution for intelligent microgrid energy management and outperforms conventional rule-based strategies and model-based optimization approaches in terms of operational cost reduction, energy utilization efficiency,...
Li Chen, Hong-Qiao Li, Zhen-Xing Chen et al.· European Conference on Elect...· 0 citations
Modern developments in electrification have rendered bidirectional Electric Vehicle (EV) charging a challenge due to the need for transactions in Vehicle-to-Grid (V2G) systems, which must address issues such as renewable generation, tariff fluctuations, and distribution grid support while also aiming to prolong device...
A. Velu, G. Naveen, S. Suraya et al.· 2026 International Conferenc...· 0 citations
M3-PPO introduces two key innovations to overcome MAML’s training instability: a Mamba-based context encoder for richer task representation in the inner loop, and a global-local momentum update mechanism for smoother meta-parameter optimization in the outer loop.
Q.-H. Dai, X. Hu, J.-L. Li et al.· Advanced Electromagnetics· 0 citations
A two-layer hierarchical control setup based on deep reinforcement learning and its main elements is looked at, and an organized forecast for the future of the field is offered, with special attention to four possible developments: multi-objective coordinated optimization, transfer learning, distributed vehicle-road co...
Hong-Bo Wang· Applied and Computational En...· 0 citations
Related blog posts
MIT News · Artificial Intelligence· news.mit.eduOct 7, 2026
Students in MIT’s Concourse program delve deeply into the human condition, debate challenging questions, and learn to develop judgment about issues that can’t be quantified.
Training AI agents with reinforcement learning can be challenging because their tools, context, and decision-making are managed by complex frameworks. Agent Lightning connects existing agents to RL training, making it easier to improve them without rebuilding them. The post Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses appeared first on Microsoft Research.
MIT News · Artificial Intelligence· news.mit.eduOct 6, 2026