Aug 2026· Proceedings of the 2026 ACM Symposium on Document Engineering· pp. 1-10· 0 citations· 31 references
TL;DR
A reference canonical digest tool that supports 13 modern and historical compressed archive formats for digest creation, validation, and similarity analysis and enables new features such as flexible integrity checks, smart file caching, remote file system attestation, and forensic similarity analysis.
Abstract
Two archives that compress the same original files will have different on-disk byte representations if they are created with different archive formats (e.g., ZIP, TAR) or compression algorithms (e.g., LZMA, DEFLATE). The ability to summarize the content and metadata of a compressed archive for later content or metadata identity analyses can help systems operate more efficiently; scan-repack-and-forward caching internet middleboxes, information processing systems, and file management and deduplication systems can all benefit from the offline ability to identify similar archives despite differences in on-disk byte representations. Many compressed archive formats are in use today, each with its own binary format, metadata fields, and often inconsistent implementations, so comparison and summarization must be flexible and format agnostic. To enable similarity comparisons between compressed archives, we propose a format-agnostic tiered canonical digest with 8 content and metadata parts. The digest's output size is independent of archive format, compression algorithm, and archive size. The digests reveal identity and similarity between archives and gracefully supports future extensions. They also enable new features such as flexible integrity checks, smart file caching, remote file system attestation, and forensic similarity analysis. We implement a reference canonical digest tool that supports 13 modern and historical compressed archive formats for digest creation, validation, and similarity analysis. Performance measurements show that digests are typically 350-440 bytes, can be computed in less than 1.5ms per KB of compressed archive, and use 300-500 KB of memory per KB of compressed archive.
Strings are the most common data type in modern database systems, yet they are often treated as an afterthought in high-performance data formats. While numerical data benefits from specialized, lightweight compression schemes, text is typically handled by general-purpose algorithms such as Zstd, LZ4, or Snappy, which r...
Tobias Schmidt, Nicolas Schmitt, Thomas Neumann et al.· Proceedings of the VLDB Endo...· 0 citations
Random access into compressed data is normally bought with density. We measure the exchange rate. Across four formats and nine axes on a common corpus, the cost of cutting a 254 MB archive into independently addressable 16 KiB units is 1.632% of the archive for an absolute-offset format against 6.57% for seekable zstd,...
With the rapid expansion of large-scale media-text corpora, multi-modal datasets increasingly require efficient storage and retrieval. Existing formats such as Files, TDP, and FFRecord work adequately for uni-modal data but expose fundamental limitations in multi-modal settings, including storage redundancy, massive sm...
A BaseX interactive XQuery environment is used to access and query a content archive of more than 5.5 million scholarly journal articles stored in a variety of XML formats. The BASEX database records metadata for each document (dual metadata if needed for different use cases), supporting fast index-based retrieval and...
Vincent Lizzi· Balisage Series on Markup Te...· 0 citations
SimPattern is a similarity detection method motivated by the characteristics of delta compression that enables more fine-grained similarity detection and substantially improves the effectiveness of delta compression, reducing the final storage size by up to 42.68% compared with state-of-the-art methods.
Zhen-Yu Cai, Xu-Ming Ye, Wen-Long Tian et al.· Proceedings of the Internati...· 0 citations