How Fast Should an Optical Circuit Switch Reconfigure for MoE Inference?
Abstract
Mixture-of-experts inference introduces fine-grained, dynamically changing many-to-many communication for token dispatch and collection, making serving latency highly sensitive to the interconnect. Optical circuit switching (OCS) offers an optically transparent data plane, but its benefits for MoE depend critically on circuit reconfiguration latency, which has not been quantified for inference workloads. We develop a performance model and an OMNeT++ simulation methodology to quantify the reconfiguration-latency budget of an OCS-based switching fabric for MoE inference. Results reveal a sharp regime transition: nanosecond-scale reconfiguration preserves favorable latency, throughput, jitter, and task completion time, whereas microsecond-scale reconfiguration collapses throughput and inflates completion time by orders of magnitude. Trace-driven replay using measured DeepSeek-V3 inference communication traces collected from an 8-H20 GPU server confirms that the same latency regimes persist under measured MoE traffic, supporting the representativeness of the model-generated workload. With Tb/s link bandwidth and increasing oversubscription, the budget tightens to the tens-of-nanoseconds regime (e.g., <inline-formula><tex-math notation="LaTeX">$44.3 \,{\mathrm{ns}}$</tex-math></inline-formula> at <inline-formula><tex-math notation="LaTeX">$1.6 \,{\mathrm{Tbps}}$</tex-math></inline-formula> and <inline-formula><tex-math notation="LaTeX">$17.2 \,{\mathrm{ns}}$</tex-math></inline-formula> at <inline-formula><tex-math notation="LaTeX">$3.2 \,{\mathrm{Tbps}}$</tex-math></inline-formula> in the evaluated configuration). Finally, we demonstrate <inline-formula><tex-math notation="LaTeX">$43.4 \,\mathrm{n}\mathrm{s}$</tex-math></inline-formula> end-to-end circuit reconfiguration on a multi-endpoint prototype, validating feasibility at the implied timescales.