Skip to content

Author

Kejiang Ye

We have 3 of 14 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Book Open access Sep 2026

ReliefServe: Relieving GPU Pressure in Multi-Model Serving via Selective CPU Escape

Sharing GPUs among many deep learning models is crucial for cost-efficient inference, but bursty multi-model workloads can easily overwhelm GPU capacity, causing severe tail latency and SLO goodput drops. Existing solutions—whether traffic-aware scheduling or hardware-level resource partitioning—can only juggle content...

Shi-Jie Peng, Yanying Lin, Cheng-Zhi Lu et al. · 0 citations
Book Open access Aug 2026

Connex: Endpoint Mobility Primitives for Dynamic LLM Serving

Evaluation on a 5-node GPU cluster under synthetic and production-derived churn shows that Connex reduces P99 tail spikes by up to 85% compared to NCCL-based baselines, achieves sub-second cutover, and maintains 100% goodput at moderate loads where baselines collapse to 0–28%, while incurring less than 5% steady-state...

Yanying Lin, Vincent Liu, Tao Luo et al. · 1 citation
Jun 2026

DynoPipe: Heterogeneous Edge-Cloud LLM Serving with Dynamically Orchestrated Pipeline Boundaries

Large language model (LLM) deployment at the network edge faces a fundamental paradox: applications require full-scale models for sophisticated reasoning, yet edge devices impose severe resource constraints across computation, memory, and network. Existing approaches fail to effectively orchestrate resources across the...

Yanying Lin, Baicheng Chen, Xinyu Zhang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.