ReliefServe: Relieving GPU Pressure in Multi-Model Serving via Selective CPU Escape
Abstract
Sharing GPUs among many deep learning models is crucial for cost-efficient inference, but bursty multi-model workloads can easily overwhelm GPU capacity, causing severe tail latency and SLO goodput drops. Existing solutions—whether traffic-aware scheduling or hardware-level resource partitioning—can only juggle contention inside the GPU pool and cannot relieve pressure when many models spike at once. We present ReliefServe, which leverages idle CPUs as interference relief paths to offload appropriate models from oversubscribed GPUs. The key is that migrating CPU-tolerant models to CPU during bursts can reduce overall GPU interference more than the local slowdown incurred, thus protecting performance-critical GPU models even with slower CPU execution. ReliefServe employs three mechanisms: (1) Escape Profiling, to identify viable offload candidates using each model’s CPU tolerance, GPU criticality, interference, and migration cost; (2) Pressure-Driven Escape Control, to monitor GPU pressure and trigger offload only if the projected interference reduction outweighs all costs; and (3) Relief-Oriented Cache Preparation, to proactively warm up high-value candidates in CPU memory for responsive migration. In a 12-server testbed with real workloads, ReliefServe reduces average inference latency by up to 27.05% and improves goodput by 39.74% under extreme burst conditions over state-of-the-art.