Preprint
Jul 2026
LMEdge: QoS-Aware LLM Inference Orchestration on Edge Clusters
This paper employs five lightweight machine learning models to predict query-specific latency, accuracy, resource usage, and response size for each model-size-quantization-device combination, and design a lightweight heuristic that approximates the BILP solution.
Reza Farahani, Zoha Azimi, Mario Colosi et al.
· 0 citations