Skip to content

Locality in Today's Content Caching: Why it Matters and How to Model it

· 0 citations · 33 references

TL;DR

A new parsimonious traffic model is proposed, named the Shot Noise Model (SNM), that enables users to natively capture the dynamics of content popularity, whilst still being suf-ficiently simple to be employed effectively for both analytical and scalable simulative studies of caching systems.

View source

Similar papers

Book Open access Aug 2026

CacheFlare: Optimizing Cold Content Performance in CDNs

Existing research on Content Delivery Networks (CDNs) predominantly focuses on optimizing the delivery of hot content—popular items that attract frequent and repeated access from large user bases. However, at Meta, we have identified that cold content, which is less popular and accessed infrequently, poses significant...

Tiansheng Zhang, YuLing Chen, Ahmed Kamal et al. · 0 citations
Open access 2026

Serverless Function Latency Model in Edge Computing

A statistical model of the service latency of serverless functions, with particular reference to edge computing, based on the observation of experimental latency values and adapted to produce a statistical distribution that closely approximates the real one is illustrated.

G. Reali, M. Femminella · 0 citations

FairCache: Demystifying Cache-Induced Unfairness in Multi-Tenant Large Language Model Serving

Large language models (LLMs) increasingly rely on context caching to enhance serving efficiency. However, this optimization inadvertently compromises fairness in multi-tenant LLM serving systems. Existing fair schedulers, which account only for compute resources, are unable to handle the multi-dimensional resource dema...

Zhuo-Yan Bai, Bin Gao, Fei Xu et al. · 0 citations
Preprint Aug 2026

CacheRoute: Planned Prefix-Affinity Routing for Large-Scale LLM Serving

When affinity recovers too little KV work, its residual load skew reduces or erases the improvement, so gating any deployment with a shadow replay rather than enabling affinity from workload statistics alone is recommended.

Huang Cheng · 1 citation
Book Open access Sep 2026

RDPart: A reuse-based OS-level cache-partitioning policy for fairness optimization in cloud data centers

RDPart is proposed, an OS-level Reuse-Driven LLC Partitioning policy designed to improve fairness while preserving the QoS of cloud workloads, and adopts a black-box design, making it well-suited for public cloud environments where real-time QoS feedback from applications is unavailable.

Javier Aznal, J. C. Saez, Carlos Bilbao · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.