Skip to content

Author

Vikram Sharma Mailthody

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Sep 2026

Fast Recovery for LLM Serving via Decoupled Device Memory Lifetime in Dynamo

Large language model (LLM) inference replicas run across tightly coupled GPUs and serve traffic continuously for weeks. Hardware and software failures are therefore inevitable, and one worker failure can disrupt an entire replica. Recovery requires reinitializing the engine, taking minutes even when weights and compila...

Schwinn Saereesitthipitak, Mohammed H. Abdulwahhab, Hannah Zhang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.