Skip to content

Author

N. Yadwadkar

We have 4 of 65 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Book Open access Sep 2026

TierSizer: Counterfactual Reasoning for DRAM Sizing in Tiered Memory Systems

Compute Express Link (CXL) enables cost-effective memory capacity expansion by placing additional tiers behind a coherent fabric. To overcome the performance loss induced by the higher latency of CXL, tiering systems keep hot data in DRAM and demote cold data to the slower tiers. While most research has focused on the...

Sujay Yadalam, Saarth Deshpande, Divyanshu Saxena et al. · 0 citations
Book Open access Sep 2026

Beyond Utilization: Energy-Conscious GPU Sharing for Inference Serving

GPUs are expensive, yet inference-serving GPU clusters remain heavily underutilized. To improve utilization, state-of-the-art systems adopt GPU multiplexing. However, optimizing solely for utilization can counterintuitively increase energy consumption. Designing policies that treat power and energy as first-order metri...

Prasoon Sinha, Dimitrios Liakopoulos, Nathan Lemma et al. · 0 citations
Jul 2026

Elastic Memory Remapping for Multi-tenant LLM Serving

This work presents Oneiros, a dynamic remapping engine for multi-tenant LLM serving that dynamically repurposes GPU memory allocated for model parameters as KV cache capacity, enabling nonblocking, unidirectional parameter transfer.

Ruihao Li, Shagnik Pal, Vineeth Narayan Pullu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.