Skip to content
Book Open access

IHS-LM: Intra-batch Hybrid Scheduling and Layer Migration for VLM Pipeline Inference Acceleration on Edge Devices

Sep 2026 · Proceedings of the International Conference on Parallel Processing · 0 citations · 39 references

TL;DR

IHS-LM, a low-latency pipeline-parallel inference framework for VLM serving on heterogeneous edge devices, introduces an Intra-Batch Hybrid Scheduling method (IHS), which dynamically adjusts the prefill–decode token ratio based on token distribution and available KV cache capacity and a Bandwidth Aware Model Layer Migration method (BAMLM), which detects imbalance caused by dynamic bandwidth and resolves imbalance through fine-grained layer migration.

Abstract

Vision-Language Models (VLMs) are increasingly deployed via distributed inference to perform visual perception and semantic reasoning tasks on edge devices. Pipeline parallelism is well-suited for collaborative edge inference, as it partitions models across devices with minimal communication between adjacent stages. However, pipeline-parallel inference on edge devices faces two challenges. First, prior intra-batch scheduling methods determine the prefill–decode token ratio without considering available KV cache capacity and token distribution, resulting in intra-batch workload imbalance. Second, static model partitioning methods assign layers to devices without accounting for dynamic bandwidth variations, leading to load imbalance across pipeline stages. These imbalances result in pipeline bubbles and increased inference latency. To address these challenges, we present IHS-LM, a low-latency pipeline-parallel inference framework for VLM serving on heterogeneous edge devices. Specifically, to alleviate intra-batch workload imbalance, IHS-LM introduces an Intra-Batch Hybrid Scheduling method (IHS), which dynamically adjusts the prefill–decode token ratio based on token distribution and available KV cache capacity. To mitigate load imbalance across stages, IHS-LM proposes a Bandwidth Aware Model Layer Migration method (BAMLM), which detects imbalance caused by dynamic bandwidth and resolves imbalance through fine-grained layer migration. We compare IHS-LM with existing pipeline-parallel inference methods in high-load and heterogeneous bandwidth settings. Experiments show that IHS-LM achieves the lowest inference latency, while preserving model accuracy.

Read PDF

Similar papers

Book Open access Aug 2026

Turbo: Efficiently Serving Long-Context Large Language Models with In-Network Aggregation

This work proposes Turbo, a first-of-its-kind in-network aggregation system that accelerates long-context inference by offloading query broadcast and attention aggregation to switches and introduces a rolling forward scheme that propagates states to enable cross-stage updates.

Ying Wan, Yuchen Xu, Chuwen Zhang et al. · 0 citations
Book Open access Sep 2026

MCSched: Memory-Controller-Aware Scheduling for Embodied LLM Workloads on NVIDIA Jetson

Preliminary but compelling evidence that MC-aware scheduling is a practical operating system/runtime direction for deployable embodied AI on unified memory edge devices is provided.

Teng Mei, Cheng-Xuan Pei, Marco Canini et al. · 1 citation
#edge computing Preprint Oct 2026

EdgeAgent: Orchestrating On-Device LLM inference for End-User Multi-Agent Systems on CPU-GPU Unified Memory Architectures

This work presents EdgeAgent, a cross-layer inference system explicitly co-designed for edge UMA and multi-agent workloads, and demonstrates that the UMA-aware execution alone contributes a 1.29x speedup over batched speculative decoding and adding the agent-aware scheduling lifts the full EdgeAgent system to a 1.77x s...

Yu-Hai Long, Yuan-Xin Wei, Kai Wu et al. · 0 citations
Book Open access Sep 2026

DASched: Dependency-Aware Scheduling for Multi-Stage Pipelines under Shared GPU Resources

Modern AI applications increasingly rely on multi-stage pipelines that link multiple computational stages into end-to-end workflows, such as video analytics, recommendation systems, and healthcare analysis. Existing serving systems optimize each stage in isolation, overlooking execution-level dependencies that arise wh...

Xue-Mei Peng, Ze-Yi Wen · 0 citations
#graph neural networks Book Open access Sep 2026

Memory-Aware Joint Optimisation of Partitioning and Scheduling for Pipeline-Parallel Training

OptPipe is presented, a unified framework that jointly optimises partitioning and scheduling for pipeline parallelism and introduces a memory-aware directed acyclic graph (DAG) that captures both task dependencies and the lifetime of intermediate tensors, enabling explicit reasoning about the trade-off between executio...

Ning Wang, A. Raith, Oliver Sinnen · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.