Skip to content
Open access

SwiftQK: Fast and Communication-Efficient Tensor Parallelism for Query-Key Normalization

Aug 2026 · IEEE computer architecture letters · pp. 1-4 · 0 citations · 12 references
Computer Science

TL;DR

SwiftQK is presented, a multi-GPU RMSNorm kernel that exchanges only scalar normalization statistics and overlaps the remaining Peer-to-Peer reduction with independent element-wise computation in a deadlock-safe persistent kernel.

Abstract

Query-Key Normalization (QK-Norm) improves the training stability and quality of modern Large Language Models (LLMs). However, under Tensor Parallelism (TP), layerwise QK-Norm introduces additional cross-GPU communication because the normalization factor depends on the full hidden vector. We present SwiftQK, a multi-GPU RMSNorm kernel that exchanges only scalar normalization statistics and overlaps the remaining Peer-to-Peer reduction with independent element-wise computation in a deadlock-safe persistent kernel. Evaluations on recent LLMs show that SwiftQK reduces QK-Norm latency by 81.4--93.9% relative to the standard TP QK-Norm using full-vector All-Gather. In end-to-end serving, SwiftQK reduces TPOT on average by 29.5% over the All-Gather-based baseline and by 14.3% over an optimized scalar-aggregation implementation.

Read PDF

Similar papers

Preprint Sep 2026

Unleashing the Power of Equality Saturation for Tensor Program Superoptimization

Efficient GPU implementations of tensor programs often require joint optimization of high-level algebraic formulations and low-level execution strategies. However, the resulting search space grows rapidly as transformations combine across operators, making joint optimization difficult to scale. We present EqiForge, a t...

Qi Zhan, Xing Hu, Xin Xia et al. · 0 citations
Open access Sep 2026

Bullseye Hash: An Efficient Hash-Table for Sparse Tensor Contraction

This work proposes Bullseye Hash, a novel hash table designed to efficiently support SpTC computations and provides guidance on configuring data object representations based on their specific characteristics, considering both algorithmic complexity and cache efficiency.

Guo-Feng Feng, Ze-Cheng Li, Ming-Zhen Li et al. · 0 citations
Preprint Aug 2026

LLMVisor: A Real-Time Latency Attribution Model for Multi-Tenant LLM Serving

LLMVisor is presented, a roofline-guided latency attribution model that captures the memory-bound and compute-bound phases via a concise piecewise-linear form over features proportional to FLOPs and memory I/O traffic and runs efficiently at microsecond scale.

Shuowei Jin, Xue-Shen Liu, Jiaxin Shan et al. · 3 citations
Preprint Aug 2026

DataKernelBench: Can LLMs Optimize Database Queries on GPUs?

DataKernelBench is introduced, which translates SQL into validated PyTorch TorchPlan programs and evaluates LLMs that optimize either the core tensor-bounded snippet or the full query in CUDA or Triton through execution-guided repair and finds that higher-performing implementations commonly use kernel fusion and execut...

Gokul Karthik Kumar, Yotam Perlitz, Corey Lammie et al. · 1 citation
#machine learning Preprint Sep 2026

GEM-KMeans: Memory-Efficient and Accurate Clustering on Massive Scale with GPU Optimization

This paper introduces GEM-KMeans, a spectrally normalized yet mathematically equivalent NLR formulation that fuses the gradient update, nonnegative projection, and sufficient statistics for normalization and iterate movement into a matrix-multiplication epilogue.

Peng Xu, Nihar Koganti, Volodymyr Kindratenko et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.