A hybrid-precision task-oriented communication framework for edge inference to holistically balance communication, on-device computation, and utility is proposed and confirmed that this design achieves an optimal trade-off among communication efficiency, on-device computational cost, and inference accuracy.
Abstract
Edge inference has emerged as a promising solution for the proliferation of artificial intelligence (AI) services by deploying models at the network edge to circumvent cloud-routing latency. Existing edge inference approaches mainly focused on either cooperative inference to reduce latency or lightweight model design to fit resource-constrained devices. These solutions often address the communication and computation challenges separately, and thus struggle to achieve a balanced trade-off among transmission efficiency, on-device processing cost, and inference accuracy. To bridge this gap, this paper proposes a hybrid-precision task-oriented communication framework for edge inference to holistically balance communication, on-device computation, and utility. In this framework, a binarized front-end is deployed on the edge device to extract and transmit binary features via orthogonal frequency-division multiplexing (OFDM) signals, while a full-precision back-end on the edge server performs the final inference. To ensure model consistency, we introduce an on-device binarization method tailored for split inference and develop an integrated channel-aware transmission scheme featuring subcarrier-based feature calibration. Furthermore, a knowledge distillation (KD)-based training strategy, supported by specialized gradient estimators, is developed to optimize the end-to-end system and inherit semantic knowledge from a full-precision teacher model. Extensive experiments on the large-scale ImageNet dataset demonstrate the superiority of the proposed hybrid system. Our analysis confirms that this design achieves an optimal trade-off among communication efficiency, on-device computational cost, and inference accuracy, outperforming existing edge inference solutions.
The advent of sixth-generation (6G) mobile networks forecasts the widespread deployment of edge artificial intelligence (AI), where AI inference tasks are offloaded from resource-constrained edge devices to edge servers. This paradigm promises low latency and high-efficiency processing, yet faces critical challenges in meeting the stringent latency requirements of emerging applications such as autonomous driving and real-time robotics. Traditional ultra-reliable and low-latency communication (URLLC) paradigms fall short in the context of edge inference, where the high-dimensional nature of extracted features introduces a fundamental trade-off between reliability and latency. In this paper, we exploit the inherent robustness of AI models to channel distortions to design an ultra-low-latency edge inference framework that jointly optimizes computation and communication resources. We focus on both multi-snapshot (sequential sensing) and multi-view (distributed sensing) scenarios. For each, we derive upper bounds on inference accuracy as a function of the number of snapshots/views and their average bit error rate (BER). These bounds guide the formulation of optimization problems that minimize total system latency while satisfying accuracy constraints. The joint optimization is decomposed into subproblems involving snapshot/view selection, transmit power control, and allocation of computation and communication time. To solve this efficiently, we develop a low-complexity iterative algorithm. Experimental results on synthetic and real-world datasets validate our approach, demonstrating significant latency reductions while maintaining inference performance. Our findings provide a foundation for robust and efficient edge AI design in future 6G networks.
Zhi-Feng Wang, Qunsong Zeng, Ren-Zhi Yuan et al.· IEEE Transactions on Wireles...· 0 citations
AceSpec, an asymmetric edge-cloud collaborative framework that employs an asymmetric communication protocol that transmits minimal main-chain indices uplink and compact sparse distributions downlink and introduces a network-aware, Lagrangian-optimized resource allocation strategy that dynamically maximizes the local cache hit rate.
Yi-Da Zhang, Zhi-Yong Gao, Shuai-Bing Yue et al.· 0 citations
DABO is proposed, a calibration-aware binary offloading method for collaborative large–small model inference that maintains competitive end-to-end accuracy while processing an average of 83.72% of requests at the edge.
Chen Zhu, Yi-Ming Su, Chenwenjie Mao et al.· IEEE Access· 0 citations
Large language models are increasingly deployed in IoT, mobile, and edge-assisted scenarios, where inference is constrained by limited computation, GPU memory capacity, heterogeneous hardware, network conditions, and cloud invocation costs. Existing selection and offloading methods mainly focus on coarse-grained model selection or execution placement, and therefore provide limited support for jointly optimizing model size, quantization precision, and deployment location. This paper proposes a task-aware model size and precision selection framework for edge-assisted LLM inference. The framework first uses an MPNet-base and ExtraTrees-based task-configuration pair router to construct a task-specific quality-feasible candidate set, and then selects the final execution configuration by considering predicted quality feasibility, model inference latency, GPU memory consumption, communication latency, and cloud invocation constraints. Experiments with multiple Qwen3 configurations show that, compared with the fixed Edge-8B FP16 baseline under the good-network setting, the proposed method improves the quality pass rate by 13.9%, reduces average latency by 17.0%, and lowers average GPU memory consumption by 33.4%. These results show that the proposed framework achieves a more favorable trade-off among answer quality, latency, and GPU memory consumption in edge-assisted LLM inference.
Yuan Yuan, Bing-Xiang Lu, Song-Wei Zhang et al.· International Conferences on...· 0 citations
The increasing scale and computational demands of large artificial intelligence models (LAIMs) present significant challenges for efficient inference in resource-constrained distributed environments. In this paper, we propose a multi-cluster LAIM co-inference framework, where an edge server equipped with multiple graphics processing units (GPUs) coordinates multiple user clusters to execute inference tasks collaboratively. Within each cluster, devices capture data from diverse perspectives and employ lightweight on-device LAIMs to extract local features. These features are then transmitted to the edge server, where they are aggregated and fused to generate a more accurate inference outcome. To reveal the fundamental trade-off between model pruning and collaborative inference performance, we develop a theoretical framework that characterizes the impact of pruning ratios and device contributions using rate-distortion theory and partial information decomposition. Based on this analysis, we formulate a joint optimization problem that determines the model pruning ratio, the task scheduling strategy, the bandwidth allocation, and the transmission power, with the goal of minimizing the inference distortion while satisfying the constraints of latency, energy consumption, and server capacity. Extensive simulation results demonstrate that the proposed framework significantly outperforms existing benchmark schemes, achieving superior inference accuracy and resource efficiency in multi-cluster edge intelligence networks.
Xiaowen Cao, Zhonghao Lyu, Shicheng Chu et al.· 0 citations
Split computing constitutes a widely used distributed inference approach, where a lightweight head model is onloaded onto the device and a heavier tail model resides on an edge server, leveraging the growing computational capabilities of modern System-on-Chips while alleviating server load. As intelligent indoor environments such as smart offices grow increasingly populated with diverse IoT devices, a single edge server must simultaneously assist multiple devices, each competing for the same shared inference resources. Without a principled mechanism to manage this shared load, the server is quickly overwhelmed, causing latency SLO violations and rendering server-assisted inference ineffective. In this work, we present MANE, a distributed inference framework that equips the server with a multi-path tail architecture, enabling a dynamic accuracy--throughput trade-off at runtime. By introducing a novel multi-path model architecture, a three-stage training scheme featuring a Joint Head Network Distillation loss and a hysteresis-based scheduler with an equitable device-fallback policy, MANE maintains over 80% SLO satisfaction rate where state-of-the-art onloading methods fail completely, while preserving accuracy 6pp higher than on-device alternatives, across up to 40 concurrent devices.
Sokratis Nikolaidis, Stylianos I. Venieris, Leonidas Malachias et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.