STATE-AWARE PREFETCHING AND CACHE POLICIES FOR LATENCY-OPTIMIZED WEB EXPERIENCES
Abstract
Modern web applications demand sustained low latency under workloads that shift across users, devices, sessions, and network conditions. Classical cache replacement policies such as Least Recently Used (LRU) and Least Frequently Used (LFU) treat every cached object identically and ignore the cost-of-miss heterogeneity that drives user-perceived latency in session-driven web workloads. This paper presents a state-aware prefetching and adaptive cache framework whose HybridScore eviction policy combines recency, frequency, latency-sensitivity, prediction confidence, bandwidth cost, and cache occupancy in a single decomposed scoring function, with a variable-order Markov predictor supplying the prediction term and a ghost-list mechanism adapting the per-signal weights online. The framework is evaluated against LRU, LFU, Adaptive Replacement Cache (ARC), and S3-FIFO across 350 trace-driven simulations spanning two workloads (175 per workload), seven cache sizes, and five random seeds. HybridScore achieves the lowest P95 and P99 latency in every cell of the experimental matrix, with non-overlapping 95% confidence intervals against the strongest baselines (ARC and LFU, both 132.3ms at 2MB). Hit-rate improvements are workload-dependent: HybridScore exceeds every baseline on the session workload and matches the strongest baselines on Zipf. Inspection of the adapted weights reveals which scoring signal was most under-weighted at initialization, an interpretability result that single-parameter self-tuning policies cannot produce.