Skip to content

Author

Animesh Trivedi

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Jul 2026

DT-RAID: A Software-Defined Tiered RAID Architecture for Heterogeneous SSDs

The rapid proliferation of cloud and AI-driven workloads has led to increasingly complex requirements for modern storage subsystems. To meet these demands, SSD controller architectures have evolved into a fragmented landscape, offering tiers of drive types optimized for endurance, performance, or capacity. More recently, SSDs have begun to differentiate regions within the same device, enabling intra-drive heterogeneity. However, integrating such heterogeneity into the existing storage stack with minimal disruption remains challenging. In this paper, we argue that storage middleware, such as RAID, is an effective control layer to address these integration challenges. We present DT-RAID, an intra-drive heterogeneity-aware RAID architecture designed for emerging SSDs. DT-RAID monitors stripe-level I/O access patterns and makes online placement decisions without requiring application modifications. It employs a lightweight heat-tracking mechanism to dynamically place frequently accessed (hot) stripes onto the higher-performance, higher-endurance tier. Using simulations based on SNIA MSR enterprise I/O traces, we demonstrate that DT-RAID improves modeled I/O performance by up to $6.8\times$ under greater tier asymmetry and extends normalized lifespan by up to $20.9\times$ compared to uniform RAID deployments.

Kun-Chi Chiang, R. Stoica, Animesh Trivedi et al. · 0 citations
#machine learning Preprint Sep 2026

Building py-kvcache: A Performance Characterization of External KV Caching for vLLM with NVMe SSDs

Prefix caching can reduce the time to first token (TTFT) of long-context LLM requests by reusing previously computed key-value (KV) states, but for short prefixes or fast GPUs, recomputation can be faster than loading from an external cache. We characterize this tradeoff in vLLM across GPU, CPU, and NVMe tiers using synthetic workloads, long-context benchmarks, production traces, and find that cache performance depends on transfer granularity, intermediate memory use, and when transfers enter the request schedule, not only on device bandwidth. These findings motivate py-kvcache, a vLLM KV Offload connector with asynchronous direct I/O, bounded shared staging, and scheduler-aware preloading, which starts disk reads while requests are still waiting, overlapping with compute. At 80k tokens, py-kvcache loading from disk is 2.0x faster than LMCache, with preloading contributing 1.34x. With GPU, CPU, and disk caching enabled, it is 1.23x faster than LMCache and within approximately 4% of the native vLLM KV Offload implementation. LongBench and SCBench show that these benefits extend to irregular prefix chains and multi-turn workloads. Bailian trace replays improve TTFT on a weaker GPU, but on an H100 the average request falls below the break-even point and GPU memory alone retains enough prefixes. External KV caching should therefore be treated as a setup specific admission decision. The py-kvcacheimplementation is available at: https://github.com/atlarge-research/py-kvcache.

Joseph Kanichai, T. De Matteis, Animesh Trivedi · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.