Skip to content
Book Open access

Canonical Digests for Compressed Archives

Aug 2026 · Proceedings of the 2026 ACM Symposium on Document Engineering · pp. 1-10 · 0 citations · 31 references

TL;DR

A reference canonical digest tool that supports 13 modern and historical compressed archive formats for digest creation, validation, and similarity analysis and enables new features such as flexible integrity checks, smart file caching, remote file system attestation, and forensic similarity analysis.

Abstract

Two archives that compress the same original files will have different on-disk byte representations if they are created with different archive formats (e.g., ZIP, TAR) or compression algorithms (e.g., LZMA, DEFLATE). The ability to summarize the content and metadata of a compressed archive for later content or metadata identity analyses can help systems operate more efficiently; scan-repack-and-forward caching internet middleboxes, information processing systems, and file management and deduplication systems can all benefit from the offline ability to identify similar archives despite differences in on-disk byte representations. Many compressed archive formats are in use today, each with its own binary format, metadata fields, and often inconsistent implementations, so comparison and summarization must be flexible and format agnostic. To enable similarity comparisons between compressed archives, we propose a format-agnostic tiered canonical digest with 8 content and metadata parts. The digest's output size is independent of archive format, compression algorithm, and archive size. The digests reveal identity and similarity between archives and gracefully supports future extensions. They also enable new features such as flexible integrity checks, smart file caching, remote file system attestation, and forensic similarity analysis. We implement a reference canonical digest tool that supports 13 modern and historical compressed archive formats for digest creation, validation, and similarity analysis. Performance measurements show that digests are typically 350-440 bytes, can be computed in less than 1.5ms per KB of compressed archive, and use 300-500 KB of memory per KB of compressed archive.

Read PDF

Similar papers

Jul 2026

Accelerating String-Heavy Queries with LLM Token Tables

Strings are the most common data type in modern database systems, yet they are often treated as an afterthought in high-performance data formats. While numerical data benefits from specialized, lightweight compression schemes, text is typically handled by general-purpose algorithms such as Zstd, LZ4, or Snappy, which r...

Tobias Schmidt, Nicolas Schmitt, Thomas Neumann et al. · 0 citations
Preprint Sep 2026

The Price of Random Access: Measuring Block Granularity Across Four Compressed Formats

Random access into compressed data is normally bought with density. We measure the exchange rate. Across four formats and nine axes on a common corpus, the cost of cutting a 254 MB archive into independently addressable 16 KiB units is 1.632% of the archive for an absolute-offset format against 6.57% for seekable zstd,...

Yakiv Shavidze · 0 citations
Preprint Aug 2026

FFSlim: An Efficient and Lightweight Format for Multi-modal Data Storage and Retrieval

With the rapid expansion of large-scale media-text corpora, multi-modal datasets increasingly require efficient storage and retrieval. Existing formats such as Files, TDP, and FFRecord work adequately for uni-modal data but expose fundamental limitations in multi-modal settings, including storage redundancy, massive sm...

Long Yang, Yu Mao, Yuhan Shao et al. · 0 citations

Unlocking an Archive of Scholarly Journal Articles with Interactive XQuery

A BaseX interactive XQuery environment is used to access and query a content archive of more than 5.5 million scholarly journal articles stored in a variety of XML formats. The BASEX database records metadata for each document (dual metadata if needed for different use cases), supporting fast index-based retrieval and...

Vincent Lizzi · 0 citations
Book Open access Sep 2026

SimPattern: A Novel Byte-Pattern Matching based Chunk Similarity Detection for Post-Deduplication

SimPattern is a similarity detection method motivated by the characteristics of delta compression that enables more fine-grained similarity detection and substantially improves the effectiveness of delta compression, reducing the final storage size by up to 42.68% compared with state-of-the-art methods.

Zhen-Yu Cai, Xu-Ming Ye, Wen-Long Tian et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.