Skip to content
Open access

REINFORCEMENT LEARNING WITH LOAD FORECASTING FOR SMART HOME ENERGY MANAGEMENT

Jul 2026 · Bulletin of Manash Kozybayev North Kazakhstan University · pp. 295-306 · 0 citations · 21 references

TL;DR

H-UPF is proposed, a hybrid intelligent framework for scalable sequential decision-making in heterogeneous environments under uncertainty that integrates probabilistic multi-horizon forecasting via a Temporal Fusion Transformer with continuous control via Proximal Policy Optimization, embedding predictive quantile distributions directly into the agent’s state representation.

Abstract

This paper proposes H-UPF (Hybrid Universal Policy with Forecasting), a hybrid intelligent framework for scalable sequential decision-making in heterogeneous environments under uncertainty. The architecture integrates probabilistic multi-horizon forecasting via a Temporal Fusion Transformer with continuous control via Proximal Policy Optimization, embedding predictive quantile distributions directly into the agent’s state representation. A Dynamic Adaptation Layer normalizes observations relative to instance-specific scales, enabling zero-shot policy transfer across environments with 18.5× variability in operating characteristics — without inter-agent communication or per-instance retraining. Validated on two real-world residential energy management datasets (REFIT: 20 UK households; CityLearn: 6 US buildings with real PV profiles), the framework achieves 88.4% of the theoretical optimum in zero-shot transfer, outperforming meta-learning (MAML-PPO) by 8.4 percentage points (Wilcoxon p = 0.003, Cohen’s d = 1.42). Ablation analysis identifies the adaptation layer as the dominant contributor (−16.2 p.p. upon removal), while probabilistic forecasting adds +6.8 p.p. through proactive scheduling. The learned policy is robust to reward parameter variations (≤3.2 p.p. sensitivity across 5× range) and supports practical deployment: 9.8 h one-time training, 18.4 ms inference per control step.

Read PDF

Similar papers

Conference Open access 2025

Machine Learning Predictive Models in Smart Home Energy Management: Progress and Challenges

: Home Energy Management Systems (HEMS) is becoming an essential part of the low-carbon economy and smart cities due to the global energy crisis and climate change issues. Conventional Home Energy Management Systems have significant difficulties in dealing with the complexity and heterogeneity of energy data, which hin...

Xuan-Ming Zhou · 0 citations
Open access Jul 2026

INTELLIGENT SCHEDULING OF PV–STORAGE–CHARGING INTEGRATED STATIONS VIA GTRXL-PPO WITH CURRICULUM LEARNING

To address the long-horizon sequential decision-making task, characterized by complex temporal dependencies, non-stationary dynamics, and high stochasticity in distribution-level PV–storage–charging systems, this paper develops a deep reinforcement learning framework that combines Gated Transformer‑XL (GTrXL) with Prox...

Song-Cheng Lu · 0 citations
Open access Aug 2026

Research on Optimal Power Grid Scheduling Based on Transfer Reinforcement Learning

M3-PPO introduces two key innovations to overcome MAML’s training instability: a Mamba-based context encoder for richer task representation in the inner loop, and a global-local momentum update mechanism for smoother meta-parameter optimization in the outer loop.

Q.-H. Dai, X. Hu, J.-L. Li et al. · 0 citations
Jul 2026

Sustainable Smart Farm Networks: A Decision Theory-Guided Deep Reinforcement Learning Approach

This work designs a decision theory (DT)-guided transfer learning (TL) framework that unifying cyber resilience and energy adaptability in agricultural monitoring, advancing methodological innovation with DT-guided TL for stable DRL convergence, and providing design insights for sustainable agricultural cyber-physical...

Dian Chen, Zelin Wan, D. Ha et al. · 0 citations
Jul 2026

DER Allocation without Load Prediction via Reinforcement Learning

A forecast-free reinforcement learning (RL) framework for DERA allocation that learns optimal policies directly from operational data, which preserves the interpretability and constraint satisfaction of DER model while adapting to stochastic demand variations through data-driven updates.

Abed AlRahman Al Makdah, Aravind Ramana, Shao-Feng Zou et al. · 0 citations
Aug 2026

Enhancing energy management in multi-zone buildings using the on-policy reinforcement learning algorithm SARSA

This work investigates a coupling-aware, decentralized formulation of the on-policy SARSA (State–Action–Reward–State–Action) algorithm for real-time HVAC control in multi-zone open-plan offices and indicates that the approach is computationally compatible with resource-constrained building energy management system (BEM...

M. A. Attia, M. A. Abdelaal, E. Sallam · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.