Sep 2026· Proceedings of the International Conference on Parallel Processing· 0 citations· 38 references
TL;DR
SimPattern is a similarity detection method motivated by the characteristics of delta compression that enables more fine-grained similarity detection and substantially improves the effectiveness of delta compression, reducing the final storage size by up to 42.68% compared with state-of-the-art methods.
Abstract
Delta compression is a crucial technique in cloud storage systems that reduces storage overhead by identifying matching byte sequences across data chunks. A key challenge in delta compression is similarity detection, i.e., determining whether two chunks are sufficiently similar to be compressed with a delta. Existing similarity detection approaches typically rely on either super-feature representations or learning-based models, both of which are highly sensitive to byte-level insertions, deletions, and local reordering, thereby reducing the efficiency of delta compression. These approaches often fail to robustly capture similarity that may manifest as unaligned blocks, fragmented regions, or reordered content. In this paper, we propose SimPattern, a similarity detection method motivated by the characteristics of delta compression. SimPattern maps each possible byte value to a unique vector through a randomly initialized byte2vec dictionary, and represents each chunk by averaging the vectors of its constituent bytes. The similarity between chunks is then evaluated using cosine similarity in the resulting vector space. To efficiently select suitable reference chunks in high-dimensional space, SimPattern employs an approximate nearest-neighbor search-based reference selection scheme. Experimental results on real-world datasets demonstrate that SimPattern enables more fine-grained similarity detection and substantially improves the effectiveness of delta compression, reducing the final storage size by up to 42.68% compared with state-of-the-art methods. Its lightweight design makes SimPattern well-suited for deployment in parallel data-reduction pipelines without degrading system throughput.
Delta compression is a promising technique for reducing the storage footprint of similar data. As the first step in delta compression, similarity detection has a significant impact on both the delta compression ratio (DCR) and the compression throughput. Unfortunately, existing approaches often make a trade-off between...
Strings are the most common data type in modern database systems, yet they are often treated as an afterthought in high-performance data formats. While numerical data benefits from specialized, lightweight compression schemes, text is typically handled by general-purpose algorithms such as Zstd, LZ4, or Snappy, which r...
Tobias Schmidt, Nicolas Schmitt, Thomas Neumann et al.· Proceedings of the VLDB Endo...· 0 citations
State-of-the-art log compressors typically rely on a decoupled “parse-then-compress” workflow, where parsing is optimized for semantic accuracy (i.e., event identification) rather than storage efficiency. Through a comprehensive empirical study, we reveal that this architectural decoupling prevents the exploitation of...
Yang Liu, Kai-Ming Zhang, Zhuang-Bin Chen et al.· Proceedings of the ACM on So...· 0 citations
A reference canonical digest tool that supports 13 modern and historical compressed archive formats for digest creation, validation, and similarity analysis and enables new features such as flexible integrity checks, smart file caching, remote file system attestation, and forensic similarity analysis.
Michael J. May, Saed Kezel· Proceedings of the 2026 ACM...· 0 citations
ByteX introduces a quantization-aware vector kernel based on SymRaBitQ, a new symmetric quantization scheme with tight theoretical guarantees that allows index construction to run directly in the quantized space accurately and efficiently without retaining a copy of full-precision vectors.
Yao Tian, Yuncheng Lu, Li-Yao Xiong et al.· 0 citations
This work introduces Matryoshka Hash Representations (MHR), a two-stage procedure that separates full-width training from prefix organization and strengthens two common pipelines: shortlisting candidates for full-precision reranking, and pruning a low-storage graph index such as LEANN.
Pei-Chun Hua, Yun-Ming Xiao· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.