Nov 2026· IEEE Transactions on Parallel and Distributed Systems· Vol 37, pp. 2423-2439· 0 citations· 38 references
Abstract
In this paper, we present a latency-aware scheduler for large-language-model (LLM) inference across mobile devices, edge servers, and a remote cloud. Our fine-grained delay model captures OFDMA uplink/downlink rates, KV-cache backhaul serialization, and profiled GPU planning-chunk resource constraints, enabling per-request routing among (i) cloud-only execution, (ii) cloud-prefill + edge-decode split execution, and (iii) edge prefill–decode mix execution. The scheduler targets high goodput and SLO attainment under hard deadlines, TTFT/TPOT targets, and guarded SM-resource constraints via a four-stage loop: windowed admission, $\delta$δ-similar prompt bucketing, tile selection, and SM-feasible prefill/decode allocation with dynamic execution adjustment. We bound padding waste, prefill-latency standard deviation, and the number of buckets, and analyze how the feasible allocation frontier is shaped by active register, SMEM, and warp walls under high-bandwidth, low-load conditions. On an A800/A6000 heterogeneous testbed, repeated real-serving experiments exercise all three routing decisions. In a Qwen3-8B heterogeneous replay, the online policy lowers mean end-to-end latency by 6.4% relative to the best fixed mean baseline and attains 100% SLO, while its P95 confidence interval overlaps those of fixed Cloud and edge-local execution. In a heterogeneous Llama-3 replay, the policy selects Cloud/Mix/Split for 18/21/6 requests; it attains 100% SLO across the four primary workload regimes and 95.6--100% across the offered-load points. Real CUDA-launch analysis further shows that the dominant Prefill/Decode kernel pair fails the joint-residency resource bound for all 37 Llama shapes on both GPUs; guarded alternative pairs account for 41.4% and 26.9% of the analyzed pair weight on A800 and A6000, respectively.
The results are packaged in the Greenfield Startup Model (GSM), which explains the priority of startups to release the product as quickly as possible, and the need to shorten time-to-market, by speeding up the development through low-precision engineering activities.
Carmine Giardino, Nicolò Paternoster, M. Unterkalmsteiner et al.· IEEE Transactions on Softwar...· 178 citations· ⚡14
Software startup companies develop innovative, software-intensive products within limited timeframes and with few resources, searching for sustainable and scalable business models.
M. Unterkalmsteiner, P. Abrahamsson, Xiaofeng Wang et al.· e-Informatica Software Engin...· 157 citations· ⚡17
This study conducts a case survey study based on the secondary data of the major pivots happened in 49 software startups, and demonstrates that customer need pivot is the most common among all pivot types.
Sohaib Shahid Bajwa, Xiaofeng Wang, Anh Nguyen-Duc et al.· Empirical Software Engineeri...· 127 citations· ⚡15
The comparison of adopter and non-adopter sample reveals three potential adoption inhibitor, security, data privacy, and portability, which underlines the importance of the technical and security perspectives for research investigating the adoption of technology.
Nattakarn Phaphoom, Xiaofeng Wang, S. Samuel et al.· Journal of Systems and Softw...· 111 citations· ⚡8
The ongoing work building a Raspberry Pi cluster consisting of 300 nodes is presented, with potential use cases being an inexpensive and green test bed for cloud computing research and a robust and mobile data center for operating in adverse environments.
P. Abrahamsson, S. Helmer, Nattakarn Phaphoom et al.· IEEE International Conferenc...· 110 citations· ⚡7
The results indicate that software developers are a slightly happy population, but the need for limiting the unhappiness of developers remains, and 219 factors representing causes of unhappiness while developing software are identified.
D. Graziotin, Fabian Fagerholm, Xiaofeng Wang et al.· International Conference on...· 84 citations· ⚡6