Skip to content
Conference

MPTAC: Memory Controller Partitioning and Traffic-Aware Contention Control in GPUs

Sep 2026 · IEEE International Conference on Application-Specific Systems, Architectures, and Processors · pp. 17-24 · 0 citations · 33 references

Abstract

The advent of cloud-based artificial intelligence and the increased digitalization of embedded systems require powerful GPUs capable of simultaneously running kernels from different software providers. To accommodate the resource isolation and Execution Time Determinism (ETD) needed with the increasing number of kernels running simultaneously, modern NVIDIA GPUs implement MIG (Multi-Instance GPU) that effectively partitions GPU resources among kernels. However, current trends show that more kernels than the available GPU partitions need to be executed simultaneously, hence, requiring that several kernels share the same GPU partition. In this work, we show that current solutions for GPU resource partitioning fail to prevent kernels running in the same partition from affecting each other's performance. We identify the request buffers in the GPU's network-on-chip and memory controller as the main sources of this contention and propose a low-overhead mechanism, MPTAC, to monitor and control it. MPTAC partitions request buffers and introduces a software-configurable threshold parameter to enable flexible Quality-of-Service (QoS), choosing between high kernel isolation when needed for critical kernels and balanced isolation performance otherwise. MPTAC is evaluated using both synthetic workloads and representative benchmarks. In the most aggressive cases, it successfully reduces contention in the GPU partition, bringing the slowdown of the analysis kernel from $3.83 \times$ to $1.43 \times$ in synthetic workloads, and from 2.35× to a more manageable 1.78× in representative benchmarks.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.