Sep 2026· Proceedings of the International Conference on Parallel Processing· 0 citations· 39 references
TL;DR
IHS-LM, a low-latency pipeline-parallel inference framework for VLM serving on heterogeneous edge devices, introduces an Intra-Batch Hybrid Scheduling method (IHS), which dynamically adjusts the prefill–decode token ratio based on token distribution and available KV cache capacity and a Bandwidth Aware Model Layer Migration method (BAMLM), which detects imbalance caused by dynamic bandwidth and resolves imbalance through fine-grained layer migration.
Abstract
Vision-Language Models (VLMs) are increasingly deployed via distributed inference to perform visual perception and semantic reasoning tasks on edge devices. Pipeline parallelism is well-suited for collaborative edge inference, as it partitions models across devices with minimal communication between adjacent stages. However, pipeline-parallel inference on edge devices faces two challenges. First, prior intra-batch scheduling methods determine the prefill–decode token ratio without considering available KV cache capacity and token distribution, resulting in intra-batch workload imbalance. Second, static model partitioning methods assign layers to devices without accounting for dynamic bandwidth variations, leading to load imbalance across pipeline stages. These imbalances result in pipeline bubbles and increased inference latency. To address these challenges, we present IHS-LM, a low-latency pipeline-parallel inference framework for VLM serving on heterogeneous edge devices. Specifically, to alleviate intra-batch workload imbalance, IHS-LM introduces an Intra-Batch Hybrid Scheduling method (IHS), which dynamically adjusts the prefill–decode token ratio based on token distribution and available KV cache capacity. To mitigate load imbalance across stages, IHS-LM proposes a Bandwidth Aware Model Layer Migration method (BAMLM), which detects imbalance caused by dynamic bandwidth and resolves imbalance through fine-grained layer migration. We compare IHS-LM with existing pipeline-parallel inference methods in high-load and heterogeneous bandwidth settings. Experiments show that IHS-LM achieves the lowest inference latency, while preserving model accuracy.
This work proposes Turbo, a first-of-its-kind in-network aggregation system that accelerates long-context inference by offloading query broadcast and attention aggregation to switches and introduces a rolling forward scheme that propagates states to enable cross-stage updates.
Ying Wan, Yuchen Xu, Chuwen Zhang et al.· Conference on Applications,...· 0 citations
Preliminary but compelling evidence that MC-aware scheduling is a practical operating system/runtime direction for deployable embodied AI on unified memory edge devices is provided.
Teng Mei, Cheng-Xuan Pei, Marco Canini et al.· Proceedings of the 17th ACM...· 1 citation
This work presents EdgeAgent, a cross-layer inference system explicitly co-designed for edge UMA and multi-agent workloads, and demonstrates that the UMA-aware execution alone contributes a 1.29x speedup over batched speculative decoding and adding the agent-aware scheduling lifts the full EdgeAgent system to a 1.77x s...
Yu-Hai Long, Yuan-Xin Wei, Kai Wu et al.· 0 citations
Modern AI applications increasingly rely on multi-stage pipelines that link multiple computational stages into end-to-end workflows, such as video analytics, recommendation systems, and healthcare analysis. Existing serving systems optimize each stage in isolation, overlooking execution-level dependencies that arise wh...
Xue-Mei Peng, Ze-Yi Wen· Workshop Proceedings of the...· 0 citations
OptPipe is presented, a unified framework that jointly optimises partitioning and scheduling for pipeline parallelism and introduces a memory-aware directed acyclic graph (DAG) that captures both task dependencies and the lifetime of intermediate tensors, enabling explicit reasoning about the trade-off between executio...
Ning Wang, A. Raith, Oliver Sinnen· Proceedings of the Internati...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.