Parallel-Aware Early-Stopping Metrics for Hyperparameter Optimization
Abstract
Parallel and distributed HPO systems evaluate many configurations concurrently, but their throughput depends heavily on when a scheduler stops weak trials and reallocates scarce CPUs, GPUs, and memory. Existing systems commonly rank live trials by validation loss, treating the early-stop metric as an implementation detail. This demo reframes metric selection as a parallel resource-scheduling problem. We present an interactive workbench for inspecting how training loss, validation loss, uncertainty-aware metrics, and stage-adaptive metrics affect Successive Halving and Hyperband decisions under different worker budgets. Across NAS-Bench-201, LCBench, and HPOBench, the underlying study shows that training loss can improve early-stage HPO outcomes by up to 24.76% over validation loss, and uncertainty-aware metrics can further improve outcomes by up to 4% under constrained budgets. The demo helps parallel-computing researchers reason about scheduler reliability, wasted worker time, and budget allocation policies for large-scale model tuning.