Skip to content
Open access

ATAC: Anchor-tail aware context parallelism for LLM training

Aug 2026 · Journal of King Saud University: Computer and Information Sciences · Vol 38 · 0 citations · 35 references

TL;DR

This work proposes ATAC, an anchor-tail aware framework for jointly optimizing packed-document construction and context-parallel sharding in large language model training and demonstrates that input-structure-aware co-design of packing and context parallelism is an effective approach to improving calibrated pipeline-level execution efficiency in large language model training.

Abstract

Large language model training commonly relies on multidimensional parallelism, including data, tensor, pipeline, and context parallelism, to support long-context and large-scale workloads. However, real pretraining corpora consist of highly heterogeneous variable-length samples, which create a complex coupling between the internal structure of packed documents and their actual execution cost. Existing approaches typically optimize upstream packing and downstream context parallelism separately, while paying limited attention to their coupled impact on execution block completion time, local load balance, and communication overhead. To address this issue, we propose ATAC, an anchor-tail aware framework for jointly optimizing packed-document construction and context-parallel sharding in large language model training. ATAC consists of two complementary components: WFAP, which constructs execution-friendly packed documents by jointly considering workload structure and sample fragmentation during packing, and ATP-CP, which performs hierarchical sharding by exploiting the anchor-tail structure of packed documents so as to mitigate intra-block stragglers while controlling additional key-value transfer overhead. In an A800-calibrated pipeline-level evaluation, ATAC improves normalized throughput over the strongest implemented baseline, achieving an overall geometric mean speedup of 1.93×\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\times $$\end{document}. Further analysis shows that WFAP substantially reduces runtime dispersion across packed documents, while ATP-CP provides a more effective balance between runtime equalization and communication overhead. Overall, this work demonstrates that input-structure-aware co-design of packing and context parallelism is an effective approach to improving calibrated pipeline-level execution efficiency in large language model training.

Read PDF

Similar papers

Jul 2026

Libra: Taming Attention Workload Skew in Long-Context LLM Training with Bounded Sequence Pool

Long-context LLM training suffers from a load-balancing problem that sequence packing does not solve. Packing samples into fixed-token sequences balances memory and linear-cost operators, but the dominant attention cost scales with the sum of squared sequence lengths. Thus, equally sized packed sequences drawn from a l...

Yan Wang, Xiu-Long Yuan, Kaiming Yang et al. · 1 citation
Book Open access Aug 2026

Turbo: Efficiently Serving Long-Context Large Language Models with In-Network Aggregation

This work proposes Turbo, a first-of-its-kind in-network aggregation system that accelerates long-context inference by offloading query broadcast and attention aggregation to switches and introduces a rolling forward scheme that propagates states to enable cross-stage updates.

Ying Wan, Yuchen Xu, Chuwen Zhang et al. · 0 citations
Conference Jul 2026

Pegasus: Accelerating Large Language Model Inference with Stateful Prefix Caching

Modern large language model (LLM) inference suffers from severe Time-To-First-Token (TTFT) bottlenecks. Existing prefix KV caching mechanisms are inherently stateless, forcing a trade-off between cross-chunk attention accuracy and online recomputation overhead. To address this issue, we propose Pegasus, a novel statefu...

Fahao Chen, Peng Li, Dongxiao Yu et al. · 0 citations
#large language models Book Open access Sep 2026

OmniPipe: Efficient, Flexible and Scalable Pipeline Parallelism for Large Model Training

Mixture-of-Experts (MoE) has become the de facto architecture for scaling large language models, offering expanded capacity with manageable compute. Pipeline parallelism (PP) is indispensable for distributed MoE training, but state-of-the-art PP schemes face three major limitations: large pipeline bubbles, high per-sta...

Jun Li, Zhi Ma, Shigang Li · 0 citations
Preprint Sep 2026

FlowTT: Exploiting Computation Flow Reuse in Irregular Tensor-Train Embedding

Tensor-Train (TT) decomposition effectively compresses large embedding tables in recommendation models, but TT-based embedding lookup remains inefficient because partially shared computation flows across input indices are not fully reused and intermediate results are repeatedly materialized off-chip between sequential...

Jongmin Seok, C. Rhee · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.