Skip to content

Automated Tensor Scheduling for Hybrid CPU-GPU LLM Inference on Consumer Devices

Jul 2026 · arXiv.org · Vol abs/2607.10183 · 0 citations · 37 references
Computer Science

TL;DR

ATSInfer combines static tensor placement with load-aware dynamic transfer and introduces asynchronous CPU-GPU coordination to efficiently schedule hardware storage, data movement, and computation across heterogeneous backends and can substantially improve the user experience of local LLM deployment on personal consumer devices.

Abstract

Running large language models on consumer devices such as laptops and desktops is challenging because model weights often exceed GPU memory capacity, making offloading inference necessary to extend effective model capacity with CPU memory. Existing offloading systems, however, typically rely on coarse layer-level or expert-level scheduling, which overlooks substantial heterogeneity among tensors within the same layer and adapts poorly to changing hardware load conditions on such devices. This paper presents ATSInfer, a hybrid CPU-GPU inference system for consumer devices that performs offloading at tensor granularity. ATSInfer combines static tensor placement with load-aware dynamic transfer, and introduces asynchronous CPU-GPU coordination to efficiently schedule hardware storage, data movement, and computation across heterogeneous backends. We implement ATSInfer and evaluate it on representative consumer platforms using both dense and MoE models. Compared with existing systems, ATSInfer improves prefill throughput by up to 1.94$\times$ and decode throughput by up to 3.29$\times$, while also increasing GPU utilization and making more effective use of PCIe bandwidth. These results show that ATSInfer can substantially improve the user experience of local LLM deployment on personal consumer devices.

View source

Similar papers

Oct 2026

SynergyScale: Optimizing Offloading and Task Partitioning for Efficient Model Training

Deep neural networks (DNNs) with billions of parameters power many important applications, but their training is fundamentally constrained by the limited on-chip memory of GPUs. This memory wall forces training to rely on distributed execution or memory offloading, both of which introduce substantial inefficiencies. Ex...

Xiaoyang Sun, Jie Xu, Zheng Wang · 0 citations
Preprint Aug 2026

DataKernelBench: Can LLMs Optimize Database Queries on GPUs?

DataKernelBench is introduced, which translates SQL into validated PyTorch TorchPlan programs and evaluates LLMs that optimize either the core tensor-bounded snippet or the full query in CUDA or Triton through execution-guided repair and finds that higher-performing implementations commonly use kernel fusion and execut...

Gokul Karthik Kumar, Yotam Perlitz, Corey Lammie et al. · 1 citation
#machine learning Open access Sep 2026

Partition-aware scheduling for mobile heterogeneous inference co-execution

This work defines the problem of partition-aware DAG scheduling for mobile heterogeneous inference, and proposes an online iterative search framework, which decomposes large DAGs into stages, focuses search on critical operators, and uses latency predictors to estimate partitioned execution without exhaustive profiling...

Zhuo-Jin Li, Marco Paolieri, L. Golubchik · 0 citations
Preprint Aug 2026

ElastiCo: Elastic Configuration and Interference-Aware Orchestration for GPU Clusters

ElastiCo is presented, an elastic co-location framework that enables training and inference workloads to safely share GPUs through three integrated mechanisms that decomposes the resulting multi-resource allocation problem into per-job configuration selection subproblems via dynamic per-resource shadow prices.

Jing-Hao Wang, Yi-Hang Zhou, Xiaoyang Sun et al. · 0 citations

K AIROX : Adaptive GPU–CPU Hybrid LLM Inference via Online Neuron Balancing

K AIROX introduces a Live Pipeline designed to prefetch neurons by predicting next-layer activation patterns, a mechanism that dynamically redistributes neurons between the GPU and CPU based on activation patterns, and a Temporal Activation Momentum cache policy to prioritize neurons with sustained utility while minimi...

Yapeng Jiang, Minghao Gan, Zi-Cong Hong et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.