Aug 2026· Journal of King Saud University: Computer and Information Sciences· Vol 38· 0 citations· 35 references
TL;DR
This work proposes ATAC, an anchor-tail aware framework for jointly optimizing packed-document construction and context-parallel sharding in large language model training and demonstrates that input-structure-aware co-design of packing and context parallelism is an effective approach to improving calibrated pipeline-level execution efficiency in large language model training.
Abstract
Large language model training commonly relies on multidimensional parallelism, including data, tensor, pipeline, and context parallelism, to support long-context and large-scale workloads. However, real pretraining corpora consist of highly heterogeneous variable-length samples, which create a complex coupling between the internal structure of packed documents and their actual execution cost. Existing approaches typically optimize upstream packing and downstream context parallelism separately, while paying limited attention to their coupled impact on execution block completion time, local load balance, and communication overhead. To address this issue, we propose ATAC, an anchor-tail aware framework for jointly optimizing packed-document construction and context-parallel sharding in large language model training. ATAC consists of two complementary components: WFAP, which constructs execution-friendly packed documents by jointly considering workload structure and sample fragmentation during packing, and ATP-CP, which performs hierarchical sharding by exploiting the anchor-tail structure of packed documents so as to mitigate intra-block stragglers while controlling additional key-value transfer overhead. In an A800-calibrated pipeline-level evaluation, ATAC improves normalized throughput over the strongest implemented baseline, achieving an overall geometric mean speedup of 1.93×\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\times $$\end{document}. Further analysis shows that WFAP substantially reduces runtime dispersion across packed documents, while ATP-CP provides a more effective balance between runtime equalization and communication overhead. Overall, this work demonstrates that input-structure-aware co-design of packing and context parallelism is an effective approach to improving calibrated pipeline-level execution efficiency in large language model training.
This paper introduces Batch- Aware Sequence Parallelism (BASP), a sequence parallelism approach that leverages batch structure to reduce communication overhead and localizing communication and improving training efficiency.
Long-context LLM training suffers from a load-balancing problem that sequence packing does not solve. Packing samples into fixed-token sequences balances memory and linear-cost operators, but the dominant attention cost scales with the sum of squared sequence lengths. Thus, equally sized packed sequences drawn from a l...
Yan Wang, Xiu-Long Yuan, Kaiming Yang et al.· arXiv.org· 1 citation
This work proposes Turbo, a first-of-its-kind in-network aggregation system that accelerates long-context inference by offloading query broadcast and attention aggregation to switches and introduces a rolling forward scheme that propagates states to enable cross-stage updates.
Ying Wan, Yuchen Xu, Chuwen Zhang et al.· Conference on Applications,...· 0 citations
Modern large language model (LLM) inference suffers from severe Time-To-First-Token (TTFT) bottlenecks. Existing prefix KV caching mechanisms are inherently stateless, forcing a trade-off between cross-chunk attention accuracy and online recomputation overhead. To address this issue, we propose Pegasus, a novel statefu...
Fahao Chen, Peng Li, Dongxiao Yu et al.· Fall Joint Computer Conferen...· 0 citations
Mixture-of-Experts (MoE) has become the de facto architecture for scaling large language models, offering expanded capacity with manageable compute. Pipeline parallelism (PP) is indispensable for distributed MoE training, but state-of-the-art PP schemes face three major limitations: large pipeline bubbles, high per-sta...
Jun Li, Zhi Ma, Shigang Li· Proceedings of the Internati...· 0 citations
Tensor-Train (TT) decomposition effectively compresses large embedding tables in recommendation models, but TT-based embedding lookup remains inefficient because partially shared computation flows across input indices are not fully reused and intermediate results are repeatedly materialized off-chip between sequential...
Jongmin Seok, C. Rhee· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.