Aug 2026· Journal of Optical Communications and Networking· Vol 18, pp. 1128-1141· 0 citations· 44 references
Abstract
The rapid growth of large-scale AI models has driven the emergence of AI data centers (AIDCs), where distributed training produces massive communication demands under multi-job concurrent execution. All-optical data center networks have emerged as a promising solution due to their high bandwidth and low latency. However, the diverse communication demands introduce severe network resource contention. To address this issue, we propose a multi-granularity collaborative scheduling (MGCS) method for distributed training workloads over an all-optical switching network architecture. It integrates optical time-slot switching (OTS) and optical circuit switching (OCS), enabling flexible allocation of network resources to accommodate diverse communication demands. MGCS follows a phase-aware layered scheduling logic. For pipeline-parallel (PP) communication, it first formulates the deterministic OTS scheduling problem as a combinatorial optimization problem and develops both a mixed-integer linear programming (MILP) model and a heuristic algorithm. It then schedules data-parallel (DP) communication through a hybrid multi-granularity scheduling method that dynamically coordinates OTS and OCS resources. We build a small-scale all-optical switching network testbed and conduct large-scale simulations to evaluate the proposed method. The results indicate that MGCS achieves up to a 30.60% reduction in epoch training time and a 49.70% reduction in the network resource occupation ratio.
Distributed optimization, learning, and inference across data centers are typical applications in the era of intelligent computing networks. These types of applications have high dependencies and are sensitive to delay. Offloading delay-sensitive dependency tasks to distributed elastic optical data center networks (EOD...
Xu Zhang, Xue Zhou, Chuan Feng et al.· Journal of Optical Communica...· 0 citations
Deploying trillion-parameter large language models across metropolitan environments is required to sustain real-time inference. Urban power constraints, however, prohibit monolithic GPU clusters, forcing the integration of distributed supernodes into a citywide compute fabric. Over 100-km distances, optical propagation...
Liang Guo, Ji-Zhuang Zhao, Wei Quan et al.· IEEE Transactions on Network...· 0 citations
Scheduling parallel data flows (a coflow) across two computation stages of a job over data center networks (DCNs) is crucial to the completion of the job. To meet the growing demands of data-intensive applications, the hybrid-switched design combining an optical circuit switch (OCS) and an electrical packet switch (EPS...
Xin Wang, Hong Shen, Hui Tian· Proceedings of the Internati...· 0 citations
The proposed MADDPG algorithm outperforms benchmark algorithms such as CO (Cloud-Only), GO (Greedy Offloading), SEP (Static Expert Partitioning), DDPG (Deep Deterministic Policy Gradient), COMA (Counterfactual Multi-Agent Policy Gradients), and MAPPO (Multi-Agent Proximal Policy Optimization) in terms of average iterat...
De-Feng Duan, Hong Liu, Li-Yun Huang et al.· Tsinghua Science and Technol...· 0 citations
QPS-ToR is proposed, which replaces NegotiaToR's scheduling logic with SW-QPS, a sliding-window algorithm originally proposed for crossbar scheduling that achieves around 90% throughput with a single low-complexity iteration.
ReTri, a bidirectional All-to-All schedule for ORNs based on the Trivance algorithm is presented, a bidirectional All-to-All schedule for ORNs based on the Trivance algorithm that improves completion time and improves reconfigurable Bruck by up to 2.1×.
Anton Juerss, Stefan Schmid· Conference on Applications,...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.