Skip to content
Book Open access

Ampel: Scheduling at the Network Cut in ML Training

Aug 2026 · Conference on Applications, Technologies, Architectures, and Protocols for Computer Communication · 0 citations · 45 references
Computer Science

Abstract

The size and communication patterns of modern ML training workloads place significant strain on datacenter fabrics. When bandwidth demand exceeds capacity, flows experience slowdowns and iteration time grows larger. Fine-grained load-balancing such as packet spraying cannot fully resolve this issue, yet it can shift the bottleneck from individual links to groups of links partitioning the network. In this work, we present Ampel: a system to schedule ML training flows at the network cuts. It leverages packet-spraying's ability to spread traffic evenly across all available paths to simplify the view of the topology into one only containing potential bottlenecks, and then bridges this new simplified model with past work on coflow scheduling. Our simulated experiments show that Ampel can reduce average training iteration time by up to 18% compared to state-of-the-art ML schedulers.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.