FNR: A Probabilistic Data Repair Scheme for Reed-Solomon Coded Storage Amid Network Fluctuations
Abstract
Reed-Solomon (RS) Code is extensively deployed in large-scale distributed storage systems to provide high data reliability. However, existing RS-based repair methods typically rely on a deterministic assumption of constant cross-rack available bandwidth, failing to account for the stochastic network dynamics inherent in production environments. Our empirical evaluations, spanning over 11 months, reveal significant throughput fluctuations, likely induced by multi-tenant contention and background traffic. Neglecting these dynamics inevitably leads to severe performance degradation and even repair failures. To address this issue, we propose Fluctuating Network Repair (FNR), a repair scheme tailored for volatile network conditions. FNR redefines the network throughput model with a probabilistic framework characterized by (cost, probability) pairs. By seamlessly integrating this model into state-of-the-art (SOTA) repair methods, FNR expands a larger search space to identify better repair solutions. Experimental results demonstrate that in soft real-time scenarios, FNR reduces repair time by 13.35%, 14.74%, and 13.89% on average compared to Random Repair (RR), AZ-Recovery (AZ), and Computation Time Priority (CTP), respectively. Furthermore, in hard real-time scenarios, FNR provides robust deadline guarantees where SOTA methods falter.