Mix-or-Split: Latency-Aware Scheduling for Edge–Cloud LLM Inference
In this paper, we present a latency-aware scheduler for large-language-model (LLM) inference across mobile devices, edge servers, and a remote cloud. Our fine-grained delay model captures OFDMA uplink/downlink rates, KV-cache backhaul serialization, and profiled GPU planning-chunk resource constraints, enabling per-req...