Skip to content

TARE: Tail Aware Evaluation of HPC Job Runtime Prediction

Jul 2026 · arXiv.org · Vol abs/2607.04935 · 0 citations · 36 references
Computer Science

TL;DR

An empirical evaluation methodology for HPC job runtime prediction that focuses on the tail, combining GeoAccuracy weighted by resource usage with decile and split analyses and UserReq achieves the highest GeoAccuracy and lowest underestimation rate on all three datasets.

Abstract

Runtime estimates affect reservation quality, backfilling opportunities, and queue delay in HPC schedulers. Under heavy tailed workloads, however, averaging over jobs can misrepresent scheduling impact because a small fraction of jobs dominates resource usage. This paper presents an empirical evaluation methodology for HPC job runtime prediction that focuses on the tail, combining GeoAccuracy weighted by resource usage with decile and split analyses. Using production traces from NREL Eagle and ALCF Mira/Intrepid, we compare XGBoost and Last2 against the user provided walltime estimate at submission (UserReq). Across all three datasets, evaluation focused on the tail changes the offline conclusion: MeanAccuracy keeps the methods relatively close, whereas GeoAccuracy reveals clearer separation and makes UserReq's strength in the upper tail visible. In the top decile, UserReq achieves the highest GeoAccuracy and lowest underestimation rate on all three datasets, and this pattern remains stable across rolling splits. We then translate this signal into a simple hybrid scheduling policy that keeps XGBoost for most jobs and routes the top decile by proxy_cost at submission to UserReq. Online replay on four production queues reduces mean wait time by up to 8% and increases backfilled jobs by 50% to 115%. These results show that offline evaluation focused on the tail better characterizes prediction quality relevant to scheduling and informs scheduling policy design.

View source

Similar papers

Preprint Aug 2026

BOOSTEDSOSA: Accelerated Inferencing for Low Variance Stochastic Online Scheduling

BOOSTEDSOSA is introduced, a dual-FPGA ML-assisted Scheduling architecture that integrates a Machine Learning predictor for expected processing times, with a novel temporal-aware training policy, enabling its use in existing HPC systems.

Adam H. Ross, Riccardo Revalor, Aryan Singh et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Leakage-Safe and Scheduler-Aware Machine Learning for Grid Job Runtime Prediction

Accurate job runtime prediction can improve scheduling-aware resource management in grid and distributed computing environments, but prediction models must be evaluated under realistic deployment constraints. This paper revisits CPU burst time prediction on the GWA-T-4 AuverGrid workload trace and reformulates it as le...

Ashfaq Ali Shafin, Khandaker Mamun Ahmed · 0 citations
Open access Sep 2026

Machine Learning-Based Runtime Prediction and Energy Optimization for HPC Job Scheduling Using the NREL Eagle Supercomputer Dataset

Accurate runtime prediction is essential for efficient HPC job scheduling, yet users chronically overesti- mate their jobs’ requirements. We analyze 7.3 million completed jobs from the NREL Eagle supercomputer and find that the problem is far worse than previously reported: median time-limit utilization is just 6.7%, w...

Renato Quispe-Vargas, Dina Maribel Yana-Yucra, Richar Andre Vilca-Solorzano et al. · 0 citations
Book Open access Sep 2026

Parallel-Aware Early-Stopping Metrics for Hyperparameter Optimization

Parallel and distributed HPO systems evaluate many configurations concurrently, but their throughput depends heavily on when a scheduler stops weak trials and reallocates scarce CPUs, GPUs, and memory. Existing systems commonly rank live trials by validation loss, treating the early-stop metric as an implementation det...

Jiawei Guan · 0 citations
#edge computing Preprint Aug 2026

PRISM: Predictive Runtime In-place Scaling and Model Selection for Edge Microservices

PRISM, a prediction-guided runtime framework that jointly selects model variants and CPU allocations for containerized edge microservices, and adapts each pipeline stage in place and minimizes predicted CPU-package energy under deadline, resource, and offline model-level Quality of Result constraints is presented.

Uwe Gropengießer, Thomas Reuter, Dominik Schön et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.