Skip to content
Conference

Tiered Speculative Decoding with Uplink Scheduling for End-Edge LLM Inference

Aug 2026 · 2026 IEEE/CIC International Conference on Communications in China (ICCC) · pp. 461-466 · 0 citations · 20 references

Abstract

With the large-scale deployment of LLM-driven services, centralized cloud inference faces increasing serving loads and latency bottlenecks. End-edge collaborative speculative decoding has emerged as a promising paradigm because its draftand-verify mechanism enables efficient collaboration between local generation and edge validation. However, existing end-edge speculative decoding schemes still suffer from uplink transmission delays and wireless scheduling uncertainty. In particular, edge-side verification is often delayed by uplink access and scheduling uncertainty, making the verification start time a major source of tail latency. Motivated by this observation, we propose a tiered cross-layer end-edge speculative decoding framework. Specifically, we decompose uplink information into critical tokens and auxiliary distributional information, so that the critical tokens alone are sufficient to trigger edgeside verification, while the auxiliary information is incorporated only if it arrives before verification completes. This improves token acceptance without blocking the critical path. We further introduce a semi-persistent uplink scheduling mechanism that reserves predictable transmission opportunities for critical-token delivery, avoiding request-grant handshakes and reducing delay accumulation across decoding steps. Experimental comparisons against baselines demonstrate that the proposed framework increases inference throughput from 35 tokens per second (TPS) to over 45 TPS while keeping the system's P95 latency stable at approximately 45 ms, thereby improving throughput and suppressing tail latency.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.