Skip to content

Author

Sukmin Cho

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Jul 2026

ELDR: Expert-Locality-Aware Decode Routing for PD-Disaggregated MoE Serving

From a request's prefill expert activations, ELDR builds an expert signature predicting the experts it will activate during generation, which reduces median TPOT by 5.9-13.9% over the strongest of four load-balancing baselines across three MoE models and two workloads.

Sang-Jun Choi, Sukmin Cho, Yifan Xiong et al. · 1 citation · ⚡1
Preprint Aug 2026

OasisKV: Scaling In-Decode KV Cache Beyond HBM with Lookahead Sparse Prefetching

OasisKV is presented, a memory-centric LLM inference system design that alleviates HBM capacity pressure by decoupling full KV-cache storage from HBM during LLM decoding and observes that future important tokens can be predicted accurately in advance using lookahead tokens drafted by speculative decoding (SD).

Can Xiao, Sukmin Cho, Junbong We et al. · 0 citations