SOS: A High-Performance Distributed Key-Value Store for Large-Scale Online Services
Large-scale online services—including web search, recommendation, and LLM inference workloads such as Retrieval-Augmented Generation (RAG) and KV-cache offloading—demand storage that handles petabyte-scale data under millisecond tail-latency SLAs. In-memory stores are cost-prohibitive at scale; disk-based systems sacri...