Robust Intelligent Computing for HPC Capacity Management: A Prediction-to-Operations Framework under Distribution Shift
Abstract
Accurate job runtime prediction is essential for effective resource management in high-performance computing (HPC) systems. However, user-provided walltime estimates are often unreliable, and the operational costs of prediction errors are asymmetric. This paper proposes a robust intelligent computing framework that connects machine learning-based runtime prediction to operational capacity management decisions, specifically addressing the challenge of distribution shift across heterogeneous HPC environments. Our methodology employs gradient-boosted models trained on submit-time metadata for prolonged-runtime risk estimation, with rigorous internal time-split testing and strict external validation using public Parallel Workloads Archive traces. To ensure reliable decision-making under distribution shift, we apply post-hoc probability recalibration, achieving well-calibrated uncertainty estimates with ECE reduced from 0.103 to 0.066. The predictive models are integrated into a discrete-event simulation framework to evaluate capacity management policies and quantify the safety--throughput trade-off. Experimental results demonstrate strong predictive performance with AUC 0.863 on an external validation trace with substantial distributional differences. The prediction-driven policy reduces overflow probability by 19.8\% compared to baseline, while systematic error-sensitivity analysis reveals how predictive uncertainty propagates into key operational metrics. These findings provide actionable insights for deploying intelligent prediction systems in real-world HPC environments.