Jul 2026· IEEE International Conference on Cloud Computing· pp. 235-245· 0 citations· 22 references
TL;DR
This work presents a novel approach that, given a generic resource-consumption model, estimates the progress of an ongoing execution in the form of a probability distribution which supports uncertainty quantification, and performs estimation via Markov-chain Monte-Carlo sampling around a maximum a posteriori estimate.
Abstract
In cloud environments, recurring workloads such as data processing jobs, CI/CD tasks and machine learning pipelines are typically executed as opaque black-box programs. Their internal progress is not directly observable, yet their resource-consumption patterns often exhibit structural similarity across runs despite variations in input data and hardware. Accurately estimating the progress of such running tasks can bring major benefits to scheduling and resource optimization. Using resource-consumption models of such tasks, derived from previous task execution monitoring, to accurately estimate a tasks cumulative progress allows detailed and accurate proactive resource steering, planning and scheduling. We present a novel approach that, given a generic resource-consumption model, estimates, based on current resource monitoring data, the progress of an ongoing execution in the form of a probability distribution which supports uncertainty quantification. The approach accounts for deviations from the model caused by varying system performance profiles, input sizes and parameters. It performs estimation via Markov-chain Monte-Carlo sampling around a maximum a posteriori estimate. Synthetic evaluations show that the approach reliably estimates task progress with an error of approximately 0.09 when suitable priors were selected. Real-world demonstrations show that the approach remains effective, and that performance degrades only slightly when the method is applied to real data. The degree of robustness to poor prior selection has also been thoroughly explored.
Experimental results on SPEC CPU benchmarks show that MPORA delivers accurate predictions under unseen inputs and distribution shifts with low overhead, while improving schedulability and response times over existing methods.
Abby Eisenklam, G. CarlosA.Montenegro, Xian Wang et al.· Euromicro Conference on Real...· 0 citations
PRISM, a prediction-guided runtime framework that jointly selects model variants and CPU allocations for containerized edge microservices, and adapts each pipeline stage in place and minimizes predicted CPU-package energy under deadline, resource, and offline model-level Quality of Result constraints is presented.
Uwe Gropengießer, Thomas Reuter, Dominik Schön et al.· 0 citations
PeakBench is a benchmark of executable multi-tool workflows with execution-grounded dependency annotations and measured resource profiles that shows that strong logical planning does not reliably translate into safe or efficient execution under resource constraints, and exposes resource information to reduce avoidable...
Zhi-Kai Chen, Xu-Xiang Zhong, Song-Yan Li et al.· 0 citations
Apache Spark is widely used for distributed data processing, but accurately predicting application execution time remains challenging because performance depends on application structure, resource configuration, and executor‐allocation behavior. This article presents two deterministic, graph‐based simulation mode...
Hina Tariq, O. Das· Software, Practice & Experie...· 0 citations
This work presents a data-driven framework that leverages historical job traces to estimate the impact of resource modifications on queue performance, and introduces the Weighted Wait-Time Score (WWS), a bounded metric that captures both typical and tail wait-time behavior.
Bipin Gaikwad, Shraddha Singh, M. Joshi et al.· Practice and Experience in A...· 0 citations
Concurrent multi-agent workflows expose future dependencies and serving-state requirements while running on heterogeneous GPU pools with time-varying load, model residency, and resource availability. The logical workflow defines the required computation, whereas its physical scheduling units, model-lifecycle actions, r...
Jing-Hao Wang, Yi-Feng Zhang, Xiao Zhou et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.