Skip to content

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Conference Aug 2026

Tiered Speculative Decoding with Uplink Scheduling for End-Edge LLM Inference

With the large-scale deployment of LLM-driven services, centralized cloud inference faces increasing serving loads and latency bottlenecks. End-edge collaborative speculative decoding has emerged as a promising paradigm because its draftand-verify mechanism enables efficient collaboration between local generation and edge validation. However, existing end-edge speculative decoding schemes still suffer from uplink transmission delays and wireless scheduling uncertainty. In particular, edge-side verification is often delayed by uplink access and scheduling uncertainty, making the verification start time a major source of tail latency. Motivated by this observation, we propose a tiered cross-layer end-edge speculative decoding framework. Specifically, we decompose uplink information into critical tokens and auxiliary distributional information, so that the critical tokens alone are sufficient to trigger edgeside verification, while the auxiliary information is incorporated only if it arrives before verification completes. This improves token acceptance without blocking the critical path. We further introduce a semi-persistent uplink scheduling mechanism that reserves predictable transmission opportunities for critical-token delivery, avoiding request-grant handshakes and reducing delay accumulation across decoding steps. Experimental comparisons against baselines demonstrate that the proposed framework increases inference throughput from 35 tokens per second (TPS) to over 45 TPS while keeping the system's P95 latency stable at approximately 45 ms, thereby improving throughput and suppressing tail latency.

Yi-Bai Liu, Yu-Xin Liang, Yu-Xin Kong et al. · 0 citations
Conference Aug 2026

LSTM-based visual sequence regression for inner-wall roughness prediction in small-diameter CFRP holes

Carbon-fiber-reinforced polymer (CFRP) aerostructures are susceptible to drilling-induced inner-wall roughness, which can degrade joint strength and fatigue reliability. Conventional contact profilometers are limited by low efficiency and potential surface damage, while the optical approach is difficult to apply to the confined, curved inner-wall of small-diameter holes. To address these challenges, this paper proposes a Long Short-Term Memory (LSTM)-based visual sequence regression method for the non-contact prediction of the arithmetic mean roughness (Ra) of CFRP drilled holes. Inner-wall images are acquired from four circumferential orientations, and a 1000-pixel grayscale sequence is extracted from the central region of each image as the sequential input. An LSTM regression network is then trained to learn the mapping between grayscale texture sequences and surface roughness. Experimental results on the test set demonstrate that the proposed method achieves an RMSE of 0.37 µm, an MAE of 0.30 µm, and an R2 of 0.93 for Ra prediction, indicating high accuracy and robust generalization. These results confirm the feasibility of non-contact quantitative roughness assessment using visual sequential features and LSTM regression, providing an effective pathway for in-process quality inspection of small-diameter CFRP holes.

You Ding, Jiahui Zeng, Freeda A. Amir et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.