The size and communication patterns of modern ML training workloads place significant strain on datacenter fabrics. When bandwidth demand exceeds capacity, flows experience slowdowns and iteration time grows larger. Fine-grained load-balancing such as packet spraying cannot fully resolve this issue, yet it can shift th...
Valerio Torsiello, Ayush Mishra, Sushovan Das et al.· Conference on Applications,...· 0 citations
Kohn-Sham density functional theory (DFT) remains the workhorse of ab initio materials simulation, yet cubic computational and quadratic memory scaling have confined calculations to a few hundred to thousands of atoms, spanning only nanometers, far below experimentally relevant length scales. We introduce XLSDFT, a lin...
Qi-Men Xu, Yu Zhang, Di-Xing Ni et al.· 0 citations
Large language models (LLMs) are increasingly used to generate, complete, and transform information in settings where their outputs can shape consequential decisions, raising concerns about their impact on demographic disparities. In this context, causal inference provides a principled basis for assessing fairness, bec...
Patrik Okanovic, T. Hoefler, Drago Plečko· 0 citations
3D Gaussian Splatting reconstructs a scene as a collection of Gaussian primitives from a set of posed photographs called the capture. The number of primitives used to represent the scene affects reconstruction quality, storage, and rendering cost. How to select this number automatically across capture scales remains un...
Afif Boudaoud, Jia-Yi Liu, A. Calotoiu et al.· 0 citations
KV cache memory and bandwidth costs grow with context length and batch size, which limits efficient long-context inference. To address this bottleneck, we introduce WUSH-KV for low-bit KV-cache quantization. It adapts WUSH, which constructs a data-aware transform from the second-order statistics of both factors in a ma...
Jiale Chen, Vage Egiazarian, Eldar Kurtic et al.· 0 citations
We introduce Maia 200, an advanced AI accelerator delivering high performance-10 145 Tflop/s FP4 and 5072 Tflop/s FP8 within a 750W TDP and 7 TB/s HBM bandwidth. Maia exemplifies a new class of Software Defined Locally Accessed Dataflow Architectures (SDLA), which explicitly program dataflow engines to orchestrate high...
Sherry Xu, M. Heddes, Jackson Peng et al.· 0 citations
A hardware-aware representation for CUTLASS kernel selection is introduced that augments candidate configurations with statically computable estimates of induced hardware behavior, showing that explicitly representing candidate-induced hardware behavior provides a useful inductive bias for learned kernel selection.
Shriram Chandran, Dominic Rinderer, Yakup Budanaz et al.· 0 citations
It is argued that utility-scale quantum architecture is primarily a cost-performance problem across a coupled quantum-classical system, leading to a blueprint for scalable quantum processing unit (QPU) design, which closely resembles the architecture of high-performance network stacks.
This work systematize the RL-for-LLM paradigm and provides a compute-centric analysis of prominent post-training algorithmic frameworks: Proximal Policy Optimization (PPO), Group Relative Policy Optimization (GRPO), as well as their variants, and develops a taxonomy of intra- and inter-model parallelism strategies for...
Maciej Besta, L. Schmidt, Lara Nonino et al.· 0 citations
Building on NCCL's device-side API, low-latency interfaces for constructing custom collective kernels are developed and used to implement new symmetric collectives in NCCL, demonstrating benefits for both AI inference and traditional HPC workloads.
Siyuan Shen, Anton Korzh, J. Bachan et al.· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.