Aug 2026· Conference on Applications, Technologies, Architectures, and Protocols for Computer Communication· 0 citations· 12 references
Computer Science
TL;DR
Weaver is a low-overhead cross-layer diagnosis framework for GPU kernels that builds a semantic execution graph from operator, kernel timeline, and warp/block evidence, distinguishes blocked and slowed kernels, and reports an interpretable root-cause chain.
Abstract
Modern GPU workloads use batching, asynchronous execution, kernel overlap, and GPU sharing to improve utilization, but these optimizations make kernel-level slowdown hard to diagnose. A target kernel may be blocked by extra synchronization or helper kernels, or slowed by concurrent kernels after it starts. Existing tools provide GPU visibility, but they either remain too heavy for continuous online use or focus on coarse-grained symptoms, leaving fine-grained kernel-level root causes to manual analysis. This poster presents Weaver, a low-overhead cross-layer diagnosis framework for GPU kernels. Weaver builds a semantic execution graph from operator, kernel timeline, and warp/block evidence, distinguishes blocked and slowed kernels, and reports an interpretable root-cause chain. Our prototype evaluation shows that Weaver continuously collects runtime evidence with low overhead and accurately localizes anomalies.
Ghost is an OS-level GPU virtualization layer integrated directly into the open-source GPU driver, using a GPU container abstraction with cgroup -like APIs for compute and memory control and privileged hardware-level scheduling and preemption for dynamic compute resource management.
In modern GPU-based non-unified heterogeneous systems, CPU-GPU communication happens via the PCI bus. Data transfers are affected by startup overhead, which underutilizes the PCI channel bandwidth for small transfers. Modern programming models, such as CUDA and OpenMP, treat each input argument to a compute kernel inde...
Dionisis-Odysseas Sotiropoulos, Sara Royuela Alcázar, Eduardo Quiñones et al.· Proceedings of the Internati...· 0 citations
This work categorizes LLM kernel operands into three inter-workgroup sharing patterns and shows that the required optimization strategies differ across categories, from simple per-workgroup pinning to subgroup-aware co-scheduling, highlighting the need for placement-aware kernel programming and smarter architectural su...
Donghyeon Joo, Sooraj Puthoor, N. Jayasena et al.· arXiv.org· 1 citation
This paper presents a zero-overhead, multi-vendor GPU compilation framework built natively into the Rust compiler (rustc) and LLVM backends, and leverages Rust's rich type system, ownership system, and strict aliasing guarantees to efficiently manage and optimize data transfers through LLVM's Offload infrastructure.
Manuel S. Drehwald, Marcelo Domínguez, Kevin Sala et al.· 1 citation
KernelGenBench is presented, the first unified multi-source and multi-chip infrastructure for evaluating LLM- and agent-generated Triton kernels and establishes operator source, hardware platform, and agentic scaffold as distinct dimensions of kernel-generation capability, and shows that success in a familiar source-ha...
Pei-Yu Zang, Jian-Hang Tao, Jia-Ling Zhang et al.· 0 citations
A per-layer, contention-aware CPU+GPU row split for transformer prefill built on a fix for MLX's lazy-graph scheduler that accelerates Llama-shaped decoder-block prefill and saves time-to-first-token on a real Qwen2.5-7B checkpoint.
Om Mohite· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.