Skip to content

M-CSN: Joint Architecture and Flow Scheduling for Metro-Scale AI Fabric Based on Supernodes

2026 · IEEE Transactions on Network Science and Engineering · Vol 13, pp. 11513-11531 · 0 citations · 78 references

Abstract

Deploying trillion-parameter large language models across metropolitan environments is required to sustain real-time inference. Urban power constraints, however, prohibit monolithic GPU clusters, forcing the integration of distributed supernodes into a citywide compute fabric. Over 100-km distances, optical propagation delays invalidate reactive congestion control for 12.8 Tbps cross-domain pipeline parallelism (PP) traffic. The resulting high bandwidth-delay product causes in-flight data to saturate edge switch buffers before the feedback loop triggers source throttling. We address this limitation with M-CSN, an AI compute architecture featuring LOTUS, a model-driven scheduling engine for metropolitan computing power networks (CPNs). LOTUS replaces delayed reactive signaling with a proactive bi-level scheduling strategy. At the macro-layer, an inverse-SLA mechanism partitions capacity among concurrent flows to guarantee performance isolation. At the micro-layer, the scheduler decouples transmission rates from dynamic window probing: it injects calculated pacing rates and static hyper-BDP windows into source RDMA queue pairs (QPs). This source-side enforcement eliminates feedback lag. Packet-level simulations demonstrate that LOTUS sustains a 92% operational load with PFC-free execution. Eliminating pause-induced jitter and buffer saturation reduces critical-flow latency by 2.6× relative to reactive baselines. This enables fragmented urban resources to operate as a unified, deterministic, supernode-based computing power network.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.