Skip to content

Author

Bing-Xiang Lu

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Conference Aug 2026

A Task-Aware Model Size and Precision Selection Framework for LLM Inference in Edge Computing

Large language models are increasingly deployed in IoT, mobile, and edge-assisted scenarios, where inference is constrained by limited computation, GPU memory capacity, heterogeneous hardware, network conditions, and cloud invocation costs. Existing selection and offloading methods mainly focus on coarse-grained model selection or execution placement, and therefore provide limited support for jointly optimizing model size, quantization precision, and deployment location. This paper proposes a task-aware model size and precision selection framework for edge-assisted LLM inference. The framework first uses an MPNet-base and ExtraTrees-based task-configuration pair router to construct a task-specific quality-feasible candidate set, and then selects the final execution configuration by considering predicted quality feasibility, model inference latency, GPU memory consumption, communication latency, and cloud invocation constraints. Experiments with multiple Qwen3 configurations show that, compared with the fixed Edge-8B FP16 baseline under the good-network setting, the proposed method improves the quality pass rate by 13.9%, reduces average latency by 17.0%, and lowers average GPU memory consumption by 33.4%. These results show that the proposed framework achieves a more favorable trade-off among answer quality, latency, and GPU memory consumption in edge-assisted LLM inference.

Yuan Yuan, Bing-Xiang Lu, Song-Wei Zhang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.