Aug 2026· Conference on Applications, Technologies, Architectures, and Protocols for Computer Communication· pp. 60-73· 0 citations· 49 references
Computer Science
TL;DR
TurboBus is presented, which pools PCIe bandwidth across co-located jobs via emerging scale-up fabrics and reduces first-token latency by up to 40% for on-demand model loading, achieves up to 1.6x throughput for KV-cache-offloaded inference, and accelerates training by up to 7%, while imposing less than 1% overhead on co-located workloads.
Abstract
GPU memory offloading is widely adopted for LLM workloads but shifts the bottleneck to GPU-CPU transfers, which can take up to 90% of the end-to-end inference/training time! Paradoxically, over 60% of PCIe bandwidth remains idle. The root cause is that PCIe links are individually bottlenecked but collectively underutilized. Bursty, phase-driven transfer patterns leave bandwidth idle both within and across jobs. We present TurboBus, which pools PCIe bandwidth across co-located jobs via emerging scale-up fabrics. TurboBus enables any GPU to borrow idle PCIe links from neighboring GPUs, even those belonging to other jobs, while preserving isolation through a privileged daemon. At the core of TurboBus, it streams data through relay GPUs with bounded memory overhead, keeps all links busy through fine-grained PCIe allocation, enables bidirectional transfers leveraging PCIe/NVLink bandwidth asymmetry, and balances fairness and completion time with a size-aware scheduling. We fully implement TurboBus and our experiments show that it reduces first-token latency by up to 40% for on-demand model loading (within 5% of the analytical optimum), achieves up to 1.6x throughput for KV-cache-offloaded inference, and accelerates training by up to 7%, while imposing less than 1% overhead on co-located workloads.
Large language models (LLMs) are increasingly deployed on edge nodes to support edge intelligence applications. To overcome limited GPU memory, offloading-based methods partition model parameters between the GPU and host memory, enabling inference on commodity hardware. However, deploying a single model instance using...
Zhen-Zheng Li, Zhi-Qing Tang, Jian-Xiong Guo et al.· IEEE Internet of Things Jour...· 0 citations
ElastiCo is presented, an elastic co-location framework that enables training and inference workloads to safely share GPUs through three integrated mechanisms that decomposes the resulting multi-resource allocation problem into per-job configuration selection subproblems via dynamic per-resource shadow prices.
Jing-Hao Wang, Yi-Hang Zhou, Xiaoyang Sun et al.· 0 citations
Ghost is an OS-level GPU virtualization layer integrated directly into the open-source GPU driver, using a GPU container abstraction with cgroup -like APIs for compute and memory control and privileged hardware-level scheduling and preemption for dynamic compute resource management.
Driven by cost and privacy constraints, many organizations deploy large language model (LLM) inference on local CPU-GPU servers. For long-context requests, the key-value (KV) cache grows linearly with sequence length and can rapidly exhaust GPU memory, reducing both the maximum supported context and request concurrency...
Jun-Wen Zhang, Wei-Ling Yang, Jian-Bin Fang et al.· Proceedings of the Internati...· 0 citations
SemaCredit is presented, a receiver controller that admits each remote-memory operation against a vector of target-resource demands and returns each component when its corresponding HBM, Atomic, or response stage completes, reducing small-operation P99 latency by 52.4% under Atomic contention and 10.2% under response i...
Fan Yang, Jiaqi Liu, Tao Jiang et al.· 0 citations
A Aegis scheduler is presented, a scheduler that adapts placement using live fabric telemetry under operator-defined contracts that reduces service p99 RPC latency, cuts SLO violations, and lowers ECN mark rate while improving utilization under production-derived workloads.
Rui Li, Shuang Cao· Conference on Applications,...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.