Skip to content
Book Open access

Execution-Driven Auto-Tuning of Static Schedules for Resource-Efficient HPC

Sep 2026 · Workshop Proceedings of the 55th International Conference on Parallel Processing · 0 citations · 10 references

Abstract

Task-based HPC runtimes are typically tuned under a performance-first requirement with the assumption that the fastest configuration requires all available cores. This paper asks whether that assumption holds, and whether high performance can be achieved with fewer resources. We study this question using SWITCHES, a task-based runtime that fixes thread-to-core mappings at compile time, making static schedule quality the determining factor and enabling offline search of valid schedules using auto-tuning tools. In this work, we present an execution-driven Cross-Entropy Method (CEM) auto-tuner that samples candidate schedules from probabilistic models (Pos and PosEdge), compiles them through the native SWITCHES workflow, and ranks them by execution time. The auto-tuner operates in two coupled modes: a performance mode that establishes the best achievable full-core schedule, and an adaptive core-saving mode that uses a learnable core-selection mask to find the smallest active-core set that preserves acceptable performance. Across SWITCHES benchmarks and dependence-intensive synthetic task graphs, the performance mode improves three of the four workloads, by up to 32.67% over the best built-in SWITCHES static policy. The core-saving mode shows that high performance rarely requires the full-core budget. A 40-task graph matches the default full-core deployment within 1.73% using only 8 of 20 hardware threads, and an 80-task graph outperforms the full-core baseline, while remaining within 5.8% of the tuned full-core optimum, using only 13 of 20 hardware threads. These results indicate that top performance and minimal CPU-resource usage are not inherently conflicting objectives, and that static schedule tuning can pursue both within a single search framework. Thus, offering practical, schedule-level control of the available resources for more efficient and sustainable HPC execution.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.