NeuroPrefetcher: Storage-Aware Sparse LLM Inference via Delta Prefetching
NeuroPrefetcher is presented, a storage-backed LLM inference system that exploits that MLP activity during autoregressive decoding has strong temporal locality, and achieves 7.9-12.0x speedup over llama.cpp across constrained memory budgets.
Nobel Dhar, Md Romyull Islam, Xuechen Zhang et al.
· 0 citations