Skip to content

Closed-Loop Decision-Focused Learning for User-Aware Cloud Orchestration under Uncertainty

Aug 2026 · 0 citations · 34 references
Computer Science

TL;DR

Heterogeneous job scheduling is formulated as a multi-objective combinatorial optimization problem (MOCOP) under uncertain constraints and a closed-loop decision-focused learning (CL-DFL) framework for cloud orchestration is proposed to improve robustness under heterogeneous workloads.

Abstract

Time-varying cloud workloads often cause resource under-utilization during off-peak periods and resource contention during peak periods. Existing prediction-then-optimization (PTO) frameworks suffer from two-stage decoupling, hindering the balance among violation rate, user satisfaction, and resource utilization. We formulate heterogeneous job scheduling as a multi-objective combinatorial optimization problem (MOCOP) under uncertain constraints and propose a closed-loop decision-focused learning (CL-DFL) framework for cloud orchestration. CL-DFL integrates a Multivariate Time-series Graph Neural Network (MTGNN)-based spatio-temporal predictor with a zeroth-order decision-focused learning (DFL) mechanism based on the tree-structured Parzen estimator (TPE). This integration establishes an end-to-end (E2E) feedback pathway between resource perception and scheduling decisions. Furthermore, we develop the GNeuro-PLS strategy by incorporating group relative policy optimization (GRPO) into cooperative local search to improve robustness under heterogeneous workloads. Extensive experiments on four real-world datasets demonstrate that CL-DFL achieves superior trade-offs among violation rate, user satisfaction, and resource utilization. It effectively controls overload risks under regular workloads and maintains resilience under highly saturated scenarios compared with state-of-the-art baselines.

View source

Similar papers

Open access 2026

A Reinforcement Learning–Driven Latency Optimization Framework for Heterogeneous Federated Learning on Edge Devices

Heterogeneous federated learning (FL) over edge networks suffers from high end-to-end latency due to coupled delays in model distribution, on-device training and upload, and server-side aggregation. Existing latency-aware FL methods typically optimize only a single stage, such as client scheduling or communication comp...

De-Sheng Sun, Li-Xing Chen, Qi-Lin Wei et al. · 0 citations
#reinforcement learning Open access Dec 2026

Optimizing Resource Allocation in Cloud Computing Environments using Reinforcement Learning

A deep reinforcement learning framework for intelligent cloud resource allocation that jointly optimizes resource utilization, Service Level Agreement compliance, infrastructure cost, and energy efficiency, and adapts to workload distribution shifts within 200 episodes without manual retuning is presented.

Msr Prasad · 0 citations
Open access Sep 2026

HRL-TaskOpt: A Hierarchical Reinforcement Learning-Based Task Scheduling Framework for Multi-Cloud and Hybrid Environments

Cloud computing has emerged as a new paradigm, which entrusts task scheduling to ensure the satisfaction of stringent constraints on latency, energy, and resources for sustainably running real-time applications. State-of-the-art natural DRL-based scheduling solutions mainly rely heavily on DRL techniques and are either...

Krishna Patwari, Raghvendra Kumar, J. Sastry · 0 citations
Open access Aug 2026

Edge-native intelligent scheduling for virtual power plants: A multi-scale perception and constrained reinforcement learning approach

EDGE-VPP is presented, an end-to-end scheduling framework that connects fine-grained load perception with safety-aware decision-making across multiple temporal scales and achieves the lowest operating cost and the fewest constraint violations among the evaluated scheduling methods.

Yuandong Jiang, Ming-Yu Ou, Jiangnan Li · 0 citations
Oct 2026

QoE-Aware Network Slicing and Resource Allocation for Stable Mobile Online Learning

Mobile online learning is challenged by the inherent conflict between dynamically evolving pedagogical demands and rigid resource provisioning infrastructures. Conventional quality of service (QoS)-driven resource allocation paradigms suffer from two critical limitations: 1) limited responsiveness to temporal-spatial f...

Ming-Zi Chen, Pei-Shun Yan, Hong-Jun Li et al. · 1 citation · ⚡1

Related blog posts

Microsoft Research Blog Jul 13, 2026

Verifying Rust cryptography in SymCrypt, from standards to code

Cryptographic code supports vital protections in modern computing systems. Learn how a new method helps verify code as developers write it while preserving speed and adaptability as it gets implemented and evolves. The post Verifying Rust cryptography in SymCrypt, from standards to code appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.