Skip to content

Author

Sai Kapil Kumar

We have 4 of 4 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#artificial intelligence Preprint Sep 2026

The Inference Engineering Pareto Atlas: Which Optimizations Dominate the Cost, Quality, and Latency Frontier?

LLM inference optimizations report speedups on different models, GPUs, prompts, and quality metrics, making them hard to compare or combine. We build a cost, quality, and latency Pareto atlas to identify the best configurations for different deployment constraints. Since exhaustive testing is impractical, we measure 54...

Srikanta Datta Tumkur, Jay Iyer, Mehar Simhadri et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Where Should the KV Cache Live? Placement Policies Across GPU, CPU, and SSD for Long-Lived Sessions

GPU high bandwidth memory is scarce and expensive, and KV caches consume much of it as chats, agent loops, and document question answering accumulate state. Systems such as Mooncake, LMCache, FlexGen, InfiniGen, and AttentionStore extend GPU memory with CPU DRAM and SSD. The harder question is which blocks belong in ea...

Srikanta Datta Tumkur, Jay Iyer, Mehar Simhadri et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Calibrate, Then Route: A Measured Study of Learned Request Routing for Disaggregated LLM Serving

A router is studied that estimates the additional completion time on each instance using exact prompt length, predicted output length, post admission KV cache pressure, and SLO class to match the goodput of round robin using six GPUs instead of seven.

Srikanta Datta Tumkur, Jay Iyer, Mehar Simhadri et al. · 0 citations
Preprint Aug 2026

More GPUs or a Smaller Cache? Tensor Parallelism versus KV Compression for Memory-Bound LLM Serving

When an LLM serving deployment runs out of KVcache room, there are two well-established ways out. Tensor parallelism shards the weights and the KV cache across two, four, or eight devices, buying memory headroom at the price of an all-reduce on every layer and a hardware bill that grows with the device count. The algor...

Srikanta Datta Tumkur, Mehar Simhadri, Anshu Bansal et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.