This work proposes HiGFRL, a Hierarchical Graph Fusion-Driven Reinforcement Learning framework, which designs a fusion-driven dual-network architecture to optimize RL decision-making and incorporates a topology-prior-guided hybrid reward mechanism that distills static topological priors into the learning process to accelerate convergence.
Abstract
Online scheduling of dependency-aware tasks in heterogeneous cloud clusters is a fundamental yet challenging problem due to the complex interplay between DAG topologies and multi-dimensional resource constraints. While DRL has shown promise, existing GNN-based approaches often struggle to efficiently model high-order topological dependencies and suffer from loose coupling between task and resource states, leading to myopic scheduling decisions. To address these limitations, we propose HiGFRL, a Hierarchical Graph Fusion-Driven Reinforcement Learning framework. HiGFRL constructs a novel three-level state representation comprising a Static Hypergraph, a Dynamic Global Graph, and a Local Bipartite Graph to explicitly model the interplay between task dependencies and real-time cluster dynamics. Specifically, we design a fusion-driven dual-network architecture to optimize RL decision-making, where a Context Fusion Allocator integrates local bipartite matching features with fused global context to execute precise task-to-node allocation, and a Global State Evaluator leverages the global dynamic graph representation to accurately estimate expected long-term cumulative reward. Furthermore, we incorporate a topology-prior-guided hybrid reward mechanism that distills static topological priors into the learning process to accelerate convergence. Extensive experiments using real-world Alibaba cluster traces demonstrate that HiGFRL significantly outperforms heuristics and DRL baselines. Specifically, in challenging large-scale high-load scenarios, HiGFRL reduces the Makespan by up to 32.55%, and optimizes the average task flow time and average task wait time by 13.58% and 13.79%, respectively. Experimental results confirm that HiGFRL not only significantly improves cluster throughput but also ensures superior QoS by substantially reducing queuing delays. Code Release:https://github.com/igeng/HiGFRL.
Dynamic cloud workflow scheduling must balance deadline satisfaction, container utilization, and energy consumption while dealing with stochastic task-execution speeds, placement-dependent communication, and coupled task and container decisions. Workflows are naturally modeled as directed acyclic graphs (DAGs), but con...
Zong-Jin Li, Shaohan Feng, Chun-Xi Yang et al.· 0 citations
Cloud computing has emerged as a new paradigm, which entrusts task scheduling to ensure the satisfaction of stringent constraints on latency, energy, and resources for sustainably running real-time applications. State-of-the-art natural DRL-based scheduling solutions mainly rely heavily on DRL techniques and are either...
Krishna Patwari, Raghvendra Kumar, J. Sastry· International Journal of Ele...· 0 citations
ReLA is an RL scheduler built on structured representation learning and aggregation that learns intra-entity representations using self-attention and convolution, captures inter-entity operation–machine interactions using cross-attention, and aggregates multi-scale representations for parallel actor-based scoring of fe...
Zheng-Yi Kwan, Wei Zhang, Aik Beng Ng et al.· Proceedings of the Internati...· 0 citations
Multi-Agent Reinforcement Learning (MARL) has emerged as a pivotal paradigm for complex decision-making in autonomous systems and air combat. While MARL has demonstrated significant potential in air combat, achieving sophisticated tactical coordination remains a non-trivial challenge. This difficulty is largely attribu...
Junlin Liu, Cheng-Wei Li, Yang Gao et al.· 0 citations
The spatial-temporal mismatch between generation and load is exacerbated by high distributed photovoltaic (PV) penetration in distribution service areas, causing power quality degradation and PV accommodation challenges. To tackle this issue, an end-to-end optimal scheduling method based on a heterogeneous graph attent...
Wei Zheng, Han Yan, Jin-Gang Qin et al.· Journal of Renewable and Sus...· 0 citations
Related blog posts
MIT News · Artificial Intelligence· news.mit.eduOct 2, 2026
Martin Trust Center Managing Director Bill Aulet introduces Dear Dreamer, a free platform for middle and high school students who want to learn about entrepreneurship.
Microsoft Research Blog· microsoft.comSep 30, 2026
Extreme space-weather events can damage power systems on Earth and degrade GPS accuracy and satellite operations. A new machine learning system can predict where damage is likely to occur 30-60 minutes before a storm arrives. The post Forecasting space weather risks on power grids appeared first on Microsoft Research.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.