Skip to content

MegaGraph: Towards Efficient Training of Large-Scale Graph Transformers with Automated Hybrid Parallelism

Sep 2026 · 0 citations · 21 references
Computer Science

TL;DR

Evaluations demonstrate that MegaGraph enables training on large-scale graphs where state-of-the-art baselines fail due to out-of-memory (OOM) errors, and achieves up to 4.51 times training speedup while maintaining model accuracy.

Abstract

Graph Transformers (GTs) offer superior representation capabilities by overcoming the depth limitations and over-smoothing issues of traditional Graph Neural Networks (GNNs). However, scaling GTs to large graphs poses critical bottlenecks. Specifically, the attention score matrix and its associated topology-aware bias matrix jointly incur significant per-layer memory overhead, and heavy graph embedding layers result in severe workload imbalances. These characteristics are unique to GT training and are not addressed by parallelism techniques designed for either conventional GNNs or Transformers, making a dedicated solution necessary. This paper introduces MegaGraph, the first automated hybrid parallelism framework designed for efficient GT training. MegaGraph designs three specialized strategies, namely graph-aware context parallelism, heterogeneous pipeline parallelism, and hybrid data parallelism, to support efficient training on large-scale graphs. However, coordinating these three parallelism strategies yields an exponentially large configuration space. To address this complexity, an automatic search engine leverages precise cost models via a Profile - Model - Search workflow to identify the optimal parallelism configuration. Evaluations demonstrate that MegaGraph enables training on large-scale graphs where state-of-the-art baselines fail due to out-of-memory (OOM) errors. The framework reduces per-device peak memory by up to 77.8\% and achieves up to 4.51$\times$ training speedup while maintaining model accuracy.

View source

Similar papers

Preprint Sep 2026

Poseidon: DAG-Guided Parallelism Search for LLM Pre-Training on Heterogeneous Clusters

Poseidon is an efficient and scalable LLM training framework designed with heterogeneity awareness, which employs two efficient, theoretically grounded strategies: stage-level pruning via early stopping with partial estimation, and layer-to-stage mapping exploiting a ridge-like distribution pattern.

Xiao-Song Chen, Shao Nie, Zhong-Min Zhao et al. · 0 citations
Sep 2026

S-GPT: Shape-aware graph partitioning and tuning for sustainable inference of large-scale deep neural networks

The rapid growth of large-scale deep neural networks has pushed inference workloads toward increasingly diverse shapes. In practical inference, these models are often invoked with tensors of varying shapes rather than a single fixed shape. Such shape variability reshapes operator loops and intermediate tensors, leading...

Hon-Gen Shao, Cen Chen, Xin Fang et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Beyond Flat Netlist: Hierarchical Graph Representation Learning for Scalable Analysis of Sequential Circuits

DeepSeq3 is introduced, a novel hierarchical framework that abstracts circuits into a two-level representation: fine-grained combinational subgraphs partitioned by flip-flops (FFs) and a high-level Super-Node Graph (SNG) that models the register-transfer structure.

Jing-Yi Zhou, Zhengyuan Shi, Jiaying Zhu et al. · 0 citations
Jul 2026

Resource-Efficient FirmCore Decomposition on Billion-Scale Multilayer Graphs

This work introduces serial and parallel algorithms for multi-core CPUs, as well as the first GPU-based algorithm for multi-core CPUs, and introduces a grid structure the authors call FC-Grid, which is exploited to distribute work among threads.

Cheng Huang, Davide Mottin, Ira Assent · 0 citations
Book Open access Sep 2026

OmniPipe: Efficient, Flexible and Scalable Pipeline Parallelism for Large Model Training

OmniPipe is proposed, a flexible bidirectional multi-pipeline parallelism scheme for unified dense and MoE LLM training that minimizes the pipeline bubble ratio while effectively overlapping EP communication with computation, enabled by the flexible and scalable parallelism scheme of bidirectional pipelines.

Jun Li, Zhi Ma, Shi-Gang Li · 0 citations
#artificial intelligence Preprint Sep 2026

QUARTET: Quad-branch cross-Attention and Random-walk Traces for Enhancing Transformers on Relational Graphs

Relational Deep Learning (RDL) models multi-table databases as heterogeneous temporal graphs, and graph transformers currently achieve state-of-the-art performance on benchmarks like RelBench. However, the current leading model, RelGT, suffers from two key limitations: its random local sampler yields loosely connected...

K. Myint, Nan Jiang, Xiang Li et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Oct 7, 2026

Discovering the value of humanistic inquiry

Students in MIT’s Concourse program delve deeply into the human condition, debate challenging questions, and learn to develop judgment about issues that can’t be quantified.

Microsoft Research Blog Oct 7, 2026

Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses

Training AI agents with reinforcement learning can be challenging because their tools, context, and decision-making are managed by complex frameworks. Agent Lightning connects existing agents to RL training, making it easier to improve them without rebuilding them. The post Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.