Skip to content
Book Open access

SimPattern: A Novel Byte-Pattern Matching based Chunk Similarity Detection for Post-Deduplication

Sep 2026 · Proceedings of the International Conference on Parallel Processing · 0 citations · 38 references

TL;DR

SimPattern is a similarity detection method motivated by the characteristics of delta compression that enables more fine-grained similarity detection and substantially improves the effectiveness of delta compression, reducing the final storage size by up to 42.68% compared with state-of-the-art methods.

Abstract

Delta compression is a crucial technique in cloud storage systems that reduces storage overhead by identifying matching byte sequences across data chunks. A key challenge in delta compression is similarity detection, i.e., determining whether two chunks are sufficiently similar to be compressed with a delta. Existing similarity detection approaches typically rely on either super-feature representations or learning-based models, both of which are highly sensitive to byte-level insertions, deletions, and local reordering, thereby reducing the efficiency of delta compression. These approaches often fail to robustly capture similarity that may manifest as unaligned blocks, fragmented regions, or reordered content. In this paper, we propose SimPattern, a similarity detection method motivated by the characteristics of delta compression. SimPattern maps each possible byte value to a unique vector through a randomly initialized byte2vec dictionary, and represents each chunk by averaging the vectors of its constituent bytes. The similarity between chunks is then evaluated using cosine similarity in the resulting vector space. To efficiently select suitable reference chunks in high-dimensional space, SimPattern employs an approximate nearest-neighbor search-based reference selection scheme. Experimental results on real-world datasets demonstrate that SimPattern enables more fine-grained similarity detection and substantially improves the effectiveness of delta compression, reducing the final storage size by up to 42.68% compared with state-of-the-art methods. Its lightweight design makes SimPattern well-suited for deployment in parallel data-reduction pipelines without degrading system throughput.

Read PDF

Similar papers

Open access Sep 2026

Efficient Similarity Detection for Delta Compression

Delta compression is a promising technique for reducing the storage footprint of similar data. As the first step in delta compression, similarity detection has a significant impact on both the delta compression ratio (DCR) and the compression throughput. Unfortunately, existing approaches often make a trade-off between...

Xiao-Zhong Jin, Hai-Kun Liu · 0 citations
Jul 2026

Accelerating String-Heavy Queries with LLM Token Tables

Strings are the most common data type in modern database systems, yet they are often treated as an afterthought in high-performance data formats. While numerical data benefits from specialized, lightweight compression schemes, text is typically handled by general-purpose algorithms such as Zstd, LZ4, or Snappy, which r...

Tobias Schmidt, Nicolas Schmitt, Thomas Neumann et al. · 0 citations
Open access Oct 2026

LogNexus: Effective Log Compression via Unified Redundancy Encoding

State-of-the-art log compressors typically rely on a decoupled “parse-then-compress” workflow, where parsing is optimized for semantic accuracy (i.e., event identification) rather than storage efficiency. Through a comprehensive empirical study, we reveal that this architectural decoupling prevents the exploitation of...

Yang Liu, Kai-Ming Zhang, Zhuang-Bin Chen et al. · 0 citations
Book Open access Aug 2026

Canonical Digests for Compressed Archives

A reference canonical digest tool that supports 13 modern and historical compressed archive formats for digest creation, validation, and similarity analysis and enables new features such as flexible integrity checks, smart file caching, remote file system attestation, and forensic similarity analysis.

Michael J. May, Saed Kezel · 0 citations
Preprint Aug 2026

ByteX: A Unified AI Search Engine at ByteDance

ByteX introduces a quantization-aware vector kernel based on SymRaBitQ, a new symmetric quantization scheme with tight theoretical guarantees that allows index construction to run directly in the quantized space accurately and efficiently without retaining a copy of full-precision vectors.

Yao Tian, Yuncheng Lu, Li-Yao Xiong et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Matryoshka Hash Representations for Model-Aware Compact Semantic Retrieval

This work introduces Matryoshka Hash Representations (MHR), a two-stage procedure that separates full-width training from prefix organization and strengthens two common pipelines: shortlisting candidates for full-precision reranking, and pruning a low-storage graph index such as LEANN.

Pei-Chun Hua, Yun-Ming Xiao · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.