2026· IEEE Transactions on Cognitive Communications and Networking· Vol 12, pp. 11277-11292· 0 citations· 49 references
Computer Science
Abstract
Optical-circuit-switched interconnects have become one of the core components for AI training due to their flexible topology reconfiguration. In contrast to the applications carried by traditional data center networks, large-scale language model training is highly sensitive to network failures, where frequent disruptions will cause gradient synchronization delays, leading to training interruptions and wasted computational resources. Existing schemes are primarily focused on specific communication patterns, without considering the fault probability distribution. As a result, unreliable links remain on critical paths. Furthermore, passive fault response mechanisms lead to inefficient topology reconfigurations, preventing network protocol convergence and making it difficult to meet the stringent stability requirements of large-scale model training. To address reliability challenges in optical-circuit-switched interconnect, we propose TopoCrafter, which leverages dual-agent deep reinforcement learning to proactively mitigate network failures. The “Topo-Agent” estimates link failure probabilities to determine reconfiguration timing and then employs a lightweight heuristic algorithm to create failure-avoidant topology that matched to traffic pattern. Concurrently, the “Route-Agent” optimizes traffic distribution. Through their strategic interaction, the agents learn holistic policies that optimally balance network reliability and communication efficiency. To improve generalization, a progressive training approach is employed, allowing the agents to adapt to complex failure environments while accelerating convergence. Under link failure scenarios, TopoCrafter maintains reliability, reducing end-to-end latency by up to 50% and maximum link utilization by approximately 20% compared to FatTree. In addition, progressive training algorithm ensures a performance degradation of less than 10% when adapting to new failure environments, and it maintains stable high performance as the network scales.
The exponential growth of Deep Neural Network (DNN) has precipitated a crisis in inter-chiplet communication, where traditional electrical interconnects struggle to meet the bandwidth density and energy efficiency requirements of massive-scale inference. While Silicon Photonics (SiPh) inherently offers the high bandwidth density and distance-independent energy efficiency required to transcend these metallic barriers, existing optical architectures remain fundamentally constrained by static topologies and prohibitive thermal reconfiguration latencies. This rigidity renders them ill-suited for the dynamic, phase-varying traffic patterns inherent in DNN workloads. To this end, we introduce APEX, a reconfigurable photonic-electrical interconnection architecture with dynamic logical topology reconfiguration algorithm engineered to resolve these scalability barriers. Central to APEX is a novel state-aware randomized greedy heuristic algorithm, which dynamically orchestrates wavelength allocation to adapt to the phase-varying traffic patterns inherent in DNN workloads. We validate the proposed approach against an optimal Integer Linear Programming (ILP) baseline, demonstrating that our linear time heuristic achieves near-optimal fidelity with a marginal energy overhead of only 5.6% in the early stages. Furthermore, it demonstrates robust scalability to synthesize full-layer network configurations where ILP solvers face combinatorial explosion. Evaluation across representative workloads reveals that APEX delivers a substantial leap in energy efficiency, achieving 0.69 pJ/bit for the BERT model—an approximate 41% reduction compared to Simba’s 1.17 pJ/bit.
Jing-Yi Chen, Mengke Ge, Hao Luo et al.· ACM Transactions on Design A...· 0 citations
Torus networks are deployed in production AI training clusters for their path diversity and low latency, but 2D Torus scales poorly: electrical packet switches compromise latency, and high-dimensional Torus introduces excessive routing complexity. We present STON (Scalable TOrus Network), a hierarchical architecture that treats a 2D Torus as a supernode and interconnects supernodes with a reconfigurable Optical Circuit Switch (OCS) for AlltoAll-dominated large-scale training networks. STON comprises three coordinated modules: (1) fragmentaware task placement, which minimizes inter-supernode traffic by reducing job fragmentation; (2) non-disruptive logical topology mapping, governed by two principles that prevent OCS reconfiguration from disrupting running tasks or partitioning multisupernode jobs; and (3) compute-phase traffic forwarding, which ensures reachability when direct OCS circuits are unavailable. STON reduces average FCT by 42.2%-61.1% across synthetic workloads and by 52.6% on a one-day Kalos production trace (under an AlltoAll traffic model for all jobs), with 95th-percentile tail latency reduced by up to 74.5%, versus a static direct-connect baseline using the same OCS hardware.
Qinwei Yang, Peirui Cao, Ruyi Zhang et al.· Fall Joint Computer Conferen...· 0 citations
To adapt to the intensive, bursty, and latency-constrained traffic from large language model (LLM) training and inference, scale-up networks are now facing tremendous challenges. Existing electrical packet switching (EPS) fabrics provide packet-level flexibility at the cost of high power consumption and long latency. Introducing optical circuit switching (OCS) in scale-up networks offers direct optical connections that can effectively reduce power consumption and latency, but OCS lacks the packet-level flexibility required by LLM inference (especially for mixture-of-experts (MoE) inference). To address these dilemmas, this work presents xSwitch, an optical-electrical-integrated (O-E-integrated) interconnect for scale-up networks, and demonstrates its effectiveness experimentally. Unlike the traditional hybrid-optical-electrical interconnects that usually place EPS and OCS in parallel, xSwitch integrates one port-count-reduced (defined by the O/E port ratio) EPS layer (ESL) on top of an OCS layer (OSL). Then, OSL can establish optical connections for xPU pairs with stable and intensive traffic demands, while bypassing the ESL for energy and latency reduction, and the dynamic and unpredictable traffic between xPUs can be provisioned by letting OSL forward it to the ESL. We prototype xSwitch with off-the-shelf components, validate its effectiveness with real-world MoE inference tasks, and also confirm its scalability with large-scale simulations. Our results indicate that for MoE inference workloads, xSwitch with a 2:1 O/E port ratio limits the average gaps to the full-EPS baseline to 1.96% in TTFT and 0.59% in TPOT.
Wei-Chi Wu, Xuanmiao Mu, Xiao-Liang Chen et al.· Conference on Applications,...· 0 citations
A novel Multi-Armed Bandit (MAB) approach is applied to optimize dynamic bandwidth allocation at the ONU layer, enabling increased user density without inducing latency burdens at the OLT and establishes a fault-resilient infrastructure suitable for next-generation converged optical networks.
K.Tara Phani, K. Kumari· Journal of optical communica...· 0 citations
Mixture-of-experts inference introduces fine-grained, dynamically changing many-to-many communication for token dispatch and collection, making serving latency highly sensitive to the interconnect. Optical circuit switching (OCS) offers an optically transparent data plane, but its benefits for MoE depend critically on circuit reconfiguration latency, which has not been quantified for inference workloads. We develop a performance model and an OMNeT++ simulation methodology to quantify the reconfiguration-latency budget of an OCS-based switching fabric for MoE inference. Results reveal a sharp regime transition: nanosecond-scale reconfiguration preserves favorable latency, throughput, jitter, and task completion time, whereas microsecond-scale reconfiguration collapses throughput and inflates completion time by orders of magnitude. Trace-driven replay using measured DeepSeek-V3 inference communication traces collected from an 8-H20 GPU server confirms that the same latency regimes persist under measured MoE traffic, supporting the representativeness of the model-generated workload. With Tb/s link bandwidth and increasing oversubscription, the budget tightens to the tens-of-nanoseconds regime (e.g., <inline-formula><tex-math notation="LaTeX">$44.3 \,{\mathrm{ns}}$</tex-math></inline-formula> at <inline-formula><tex-math notation="LaTeX">$1.6 \,{\mathrm{Tbps}}$</tex-math></inline-formula> and <inline-formula><tex-math notation="LaTeX">$17.2 \,{\mathrm{ns}}$</tex-math></inline-formula> at <inline-formula><tex-math notation="LaTeX">$3.2 \,{\mathrm{Tbps}}$</tex-math></inline-formula> in the evaluated configuration). Finally, we demonstrate <inline-formula><tex-math notation="LaTeX">$43.4 \,\mathrm{n}\mathrm{s}$</tex-math></inline-formula> end-to-end circuit reconfiguration on a multi-endpoint prototype, validating feasibility at the implied timescales.
Shuo Li, Hua-Xi Gu, Yi-Xuan Hao et al.· Journal of Lightwave Technol...· 0 citations
ReTri, a bidirectional All-to-All schedule for ORNs based on the Trivance algorithm is presented, a bidirectional All-to-All schedule for ORNs based on the Trivance algorithm that improves completion time and improves reconfigurable Bruck by up to 2.1×.
Anton Juerss, Stefan Schmid· Conference on Applications,...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.