IHS-LM: Intra-batch Hybrid Scheduling and Layer Migration for VLM Pipeline Inference Acceleration on Edge Devices
IHS-LM, a low-latency pipeline-parallel inference framework for VLM serving on heterogeneous edge devices, introduces an Intra-Batch Hybrid Scheduling method (IHS), which dynamically adjusts the prefill–decode token ratio based on token distribution and available KV cache capacity and a Bandwidth Aware Model Layer Migr...