Skip to content
Preprint

FlashBoot: Sub-Second Weight Loading for Large Models at Rack Scale

Aug 2026 · 0 citations · 19 references
Computer Science

TL;DR

FlashBoot is presented, a hardware-friendly, framework-workflow co-designed weight-loading subsystem built on SGLang that accelerates single-node weight loading by up to 50x and concurrent rack-level weight loading by>270x and scales poorly to concurrent multi-node bring-up.

Abstract

Flagship Mixture-of-Experts (MoE) models are growing fast along two axes at once: total parameter count and the number of experts. In elastic deployment scenarios, many GPUs across many nodes must become serving-ready quickly, and this growth makes weight loading a noticeable part of the latency budget. Even on NVIDIA's GB300 NVL72, today's state-of-the-art loaders leave most of that bandwidth unused. The losses are structural: (C1) weight memory is fragmented into tens of thousands of per-tensor objects, so transfers run far below link bandwidth; (C2) cross-node replication is gated by NCCL communicator setup, which costs 10-110 s before a single weight byte moves; and (C3) the existing cross-node GPU->GPU clone path is serial and scales poorly to concurrent multi-node bring-up. We present FlashBoot, a hardware-friendly, framework-workflow co-designed weight-loading subsystem built on SGLang. At its core is FabricArena, a contiguous, exportable and inter-node addressable tensor memory layout. On top of it, FlashLoad loads from CPU as a single bulk, zero-copy transfer, and FlashClone replicates a resident model from a remote GPU via a remote-mapping mechanism that removes NCCL setup. In experiments on NVL72 with DeepSeek-V4-Pro and DeepSeek-V4-Flash, FlashClone maps remote weight memory in ~10 ms (versus 10-110 s for NCCL) and sustains>=700 GB/s per clone. Against the state of the art, FlashBoot accelerates single-node weight loading by up to 50x (from 20.1 s to 0.4 s) and concurrent rack-level weight loading by>270x (from 87 s to 0.32 s). Our code will be made publicly available.

View source

Similar papers

Preprint Aug 2026

LLMVisor: A Real-Time Latency Attribution Model for Multi-Tenant LLM Serving

LLMVisor is presented, a roofline-guided latency attribution model that captures the memory-bound and compute-bound phases via a concise piecewise-linear form over features proportional to FLOPs and memory I/O traffic and runs efficiently at microsecond scale.

Shuowei Jin, Xue-Shen Liu, Jiaxin Shan et al. · 3 citations
Preprint Aug 2026

Completion-Path Credits: Multi-Resource Control for Scale-Up Fabrics

SemaCredit is presented, a receiver controller that admits each remote-memory operation against a vector of target-resource demands and returns each component when its corresponding HBM, Atomic, or response stage completes, reducing small-operation P99 latency by 52.4% under Atomic contention and 10.2% under response i...

Fan Yang, Jiaqi Liu, Tao Jiang et al. · 0 citations
Preprint Aug 2026

The Ingestion Tax: Adopting File-Backed Weights in Tensor Frameworks

This work presents file-backed weight adoption: a framework-independent producer maps each tensor with MAP_SHARED, wraps the pages as a no-copy GPU buffer, and exports a DLPack capsule that PyTorch or MLX imports as ordinary storage.

Yuan Si, Yufeng Lin, Da-Ming Li et al. · 1 citation
Open access Aug 2026

Calibration-free compression brings Evo 2 to its full million-token context on a single GPU

TurboQuant-Bio, an open toolkit that compresses Evo 2’s weights and attention cache to four bits without calibration data, and serves both through fused kernels, proves near-lossless across perplexity spanning the tree of life, genomic classification, splice-site prediction, gene completion and clinically relevant vari...

Michail Patsakis, Alexandros Tzanakakis, I. Georgakopoulos-Soares · 0 citations
Preprint Aug 2026

Every Expert Counts: ExactMoE for Memory-Efficient W4A16 Inference

ExactMoE, an inference design that applies symmetric group-128 four-bit weight quantization only to routed experts, stores those experts in kernel-native MARLIN form in pinned host memory, and executes all selected experts through a configurable GPU-resident slot cache and fused grouped MoE kernels, identifies a practi...

Amjad Saab · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.