We have 2 of 3 papers
We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.
Not the right person? Other researchers publish under this name.
IHS-LM: Intra-batch Hybrid Scheduling and Layer Migration for VLM Pipeline Inference Acceleration on Edge Devices
IHS-LM, a low-latency pipeline-parallel inference framework for VLM serving on heterogeneous edge devices, introduces an Intra-Batch Hybrid Scheduling method (IHS), which dynamically adjusts the prefill–decode token ratio based on token distribution and available KV cache capacity and a Bandwidth Aware Model Layer Migr...