Review
Aug 2026
When Does Distributed AI Inference Need More Wide-Area Bandwidth? A Co-Design Evaluation of Optical, Packet, and Software Levers
A workload model predicting when moving inference state across sites beats recomputing it is derived, and five sensitivity axes are quantify: context length, attention architecture, queueing, agentic compounding, and loss/jitter-induced bandwidth collapse are quantified.
C. Prasanna
· 0 citations