With the large-scale deployment of LLM-driven services, centralized cloud inference faces increasing serving loads and latency bottlenecks. End-edge collaborative speculative decoding has emerged as a promising paradigm because its draftand-verify mechanism enables efficient collaboration between local generation and edge validation. However, existing end-edge speculative decoding schemes still suffer from uplink transmission delays and wireless scheduling uncertainty. In particular, edge-side verification is often delayed by uplink access and scheduling uncertainty, making the verification start time a major source of tail latency. Motivated by this observation, we propose a tiered cross-layer end-edge speculative decoding framework. Specifically, we decompose uplink information into critical tokens and auxiliary distributional information, so that the critical tokens alone are sufficient to trigger edgeside verification, while the auxiliary information is incorporated only if it arrives before verification completes. This improves token acceptance without blocking the critical path. We further introduce a semi-persistent uplink scheduling mechanism that reserves predictable transmission opportunities for critical-token delivery, avoiding request-grant handshakes and reducing delay accumulation across decoding steps. Experimental comparisons against baselines demonstrate that the proposed framework increases inference throughput from 35 tokens per second (TPS) to over 45 TPS while keeping the system's P95 latency stable at approximately 45 ms, thereby improving throughput and suppressing tail latency.
Yi-Bai Liu, Yu-Xin Liang, Yu-Xin Kong et al.· 2026 IEEE/CIC International...· 0 citations
Artificial Intelligence-Generated Content is reshaping interactive content creation. However, provisioning diffusion-based interactive denoising services in mobile edge networks remains challenging, due to the demanding computation and communication resources. To address these challenges, in this paper, we investigate the innate features of latent diffusion models, which have been widely used for image generation. We first observe that, 1) latent states are compact in data size compared to generated images, and 2) image attributes can be differentiated in frequency domain across the denoising stages. Motivated by these insights, we propose an end-edge collaborative denoising system to accelerate progressive image generation through hybrid latent state reuse and denoising tasks offloading. At the core of this system design are two synergistic algorithms configuring generation quality and service latency. The Adaptive Reuse Configuration Selection algorithm properly configures the synthesis of latent states by identifying reusable latent states, while preserving consistent image attributes between progressive tasks, thus minimizing computational redundancy with guaranteed generation quality. Subsequently, the Distributed Denoising Management algorithm optimizes the offloading ratio between the edge server and mobile users based on their computing resource availability. This paradigm enables the transmission of lightweight latent tensors instead of high-resolution images, thereby alleviating transmission burden. Extensive experimental comparisons against baselines demonstrate that, the proposed system can reduce service latency by up to 66.9% and enhance generation quality by up to 62.3%.
Yu-Xin Liang, Peng Yang, Zi-Qi Zhou et al.· IEEE Transactions on Mobile...· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.