BOOSTEDSOSA is introduced, a dual-FPGA ML-assisted Scheduling architecture that integrates a Machine Learning predictor for expected processing times, with a novel temporal-aware training policy, enabling its use in existing HPC systems.
Abstract
Heterogeneous scheduling in stochastic, online envi- ronments, such as high-performance computing (HPC) systems, presents a significant challenge. Stochastic Online Scheduling Accelerators (SOSAs) offer a promising solution, but their effectiveness is compromised by a reliance on runtime estimates provided by users. These estimates introduce substantial vari- ance into the scheduling process (mean MAE in hundreds of Core-Days), thereby weakening the competitiveness of Stochastic Online Scheduling algorithms as their competitive-ratio bound increases with runtime variability. To address this limitation, we introduce BOOSTEDSOSA, a dual-FPGA ML-assisted Scheduling architecture that integrates a Machine Learning predictor for expected processing times, with a novel temporal-aware training policy. The predictor estimates job runtimes using only scheduler parameters available at submission time, enabling its use in existing HPC systems. Using historical real-world HPC job data (from the Argonne Leadership Comput- ing Facility, MIT Supercloud and UIUC Blue Waters workload datasets), we show that the predictor reduces MAE by up to 63.85% compared to user runtime estimates, and the additive training policy reduces MAE by up to 71.88% compared to a static model. End-to-end, BOOSTEDSOSA achieves an average 17x speedup over an AVX-optimized software baseline and processes up to 1,711 jobs/seconds
Accurate runtime prediction is essential for efficient HPC job scheduling, yet users chronically overesti- mate their jobs’ requirements. We analyze 7.3 million completed jobs from the NREL Eagle supercomputer and find that the problem is far worse than previously reported: median time-limit utilization is just 6.7%, w...
Renato Quispe-Vargas, Dina Maribel Yana-Yucra, Richar Andre Vilca-Solorzano et al.· Informatica· 0 citations
This work presents a data-driven framework that leverages historical job traces to estimate the impact of resource modifications on queue performance, and introduces the Weighted Wait-Time Score (WWS), a bounded metric that captures both typical and tail wait-time behavior.
Bipin Gaikwad, Shraddha Singh, M. Joshi et al.· Practice and Experience in A...· 0 citations
Python is widely used in scientific research because it enables rapid development and provides rich ecosystems for data analysis, artificial intelligence (AI), and machine learning. However, customized research code can become prohibitively slow as experiments scale. This challenge is particularly acute in discrete-eve...
Current GPU Collective Communication Libraries (CCLs) employ predefined schedules optimized for stable environments. Their supported schedules and selection logic are fixed at communicator initialization, which fails to account for evolving runtime conditions, such as workload characteristics and hardware health status...
Rui Ding, Xiandong Lu, Jiajun Wang et al.· Conference on Applications,...· 0 citations
Deploying heterogeneous AI models concurrently on a shared GPU introduces resource contention that complicates runtime scheduling. While surrogate models avoid costly online benchmarking, their profiling requirements typically grow combinatorially with the number of co-running models, limiting scalability. We propose a...
Concord, a novel GPU sharing-enabled workload scheduler that outperforms state-of-the-art schedulers, achieves a 1.68 × reduction in JCT and a 29% improvement in GPU utilization in high-load scenarios.
Xin-Hua Wang, Wei-Wei Lin, Hai-Jie Wu et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.