Jul 2026· IEEE International Symposium on High-Performance Parallel Distributed Computing· pp. 207-220· 0 citations· 53 references
Computer Science
TL;DR
This paper develops a novel framework that characterizes compressibility limits for scientific datasets under realistic tiling constraints, and is the first framework to rigorously characterize lossy compressibility limits for scientific datasets and compressor, moving beyond classical asymptotic 1D source models.
Abstract
Error-bounded lossy compressors have been developed for years to reduce the vast volumes of scientific data generated by high-performance computing (HPC) applications and advanced scientific instruments. While these compressors have been effective in mitigating the challenges posed by massive datasets, a significant gap remains in our understanding of the fundamental compressibility limits of scientific data–an issue that critically impacts the sustainable adoption and development of efficient lossy compression techniques in practice. Classical rate-distortion theory, established by Shannon, assumes stationary 1D sources with unconstrained coding–assumptions that do not hold for scientific datasets compressed under the tiling constraints imposed by modern parallel lossy compressors. This paper addresses this gap by developing a novel framework that characterizes compressibility limits for scientific datasets under realistic tiling constraints. The contribution is two-fold. First, we establish a tile-aware, finite-blocklength extension of rate–distortion theory that advances classical 1D asymptotic formulations into a rigorous framework for piecewise 2D Gaussian random fields. To our knowledge, this is the first framework to rigorously characterize lossy compressibility limits for scientific datasets and compressor, moving beyond classical asymptotic 1D source models. Second, we conduct a comprehensive validation of the proposed modeling framework using state-of-the-art error-bounded lossy compressors and diverse real-world HPC datasets, demonstrating that our theory accurately predicts rate-distortion trends and provides actionable insights for compressor design.
With the rapid advancement of large-scale scientific simulations, the massive volume of point cloud data generated has increasingly become a critical bottleneck for scientific storage systems and data management pipelines. Existing point cloud compression techniques integrated into scientific storage systems are designed for sparse geometry and rely on quantization schemes whose optimality assumptions do not hold for dense data. When applied at the compression layer to point clouds, this representation mismatch leads to fundamentally sub-optimal rate-distortion trade-offs that cannot be addressed through parameter tuning or framework-level adaptations. This mismatch increases storage overhead and limits efficient movement and downstream analysis of simulation outputs. This issue arises in scientific data management workflows handling large-scale dense particle datasets. State-of-the-art compression methods fail to fully exploit the redundancies inherent in such data.
We address this limitation by developing a theory of point cloud compressibility for dense data, characterizing fundamental ratedistortion behavior at the representation layer. Guided by this analysis, we introduce XnYZip, an error-bounded lossy compressor based on provably optimal Truncated Octahedron quantization, combined with a locality-aware encoding pipeline using space-filling curves and run-length encoding. Experiments on large-scale scientific datasets demonstrate consistent storage and throughput improvements, achieving up to 3× higher compression ratios, 2.2× faster compression, and 1.2× faster decompression compared to state-of-the-art point cloud compressors under same distortion.
You-Yuan Liu, Longtao Zhang, Ruoyu Li et al.· Proceedings of the VLDB Endo...· 0 citations
Cortado is compared with state-of-the-art approaches and it is found that it is the only progressive compressor that guarantees the evaluated error bounds on all tested inputs while achieving the highest throughputs, comparable reconstruction quality, and on-par compression ratios.
Error-bounded lossy compression is essential for storing and transferring the vector-field data produced by large-scale scientific simulations. Although it enforces a user-specified error bound to limit numerical distortion, it does not preserve the field's topology: small admissible perturbations can create or eliminate critical points on which downstream feature analysis depends. Existing GPU compressors achieve high throughput but are topology-agnostic, whereas the only compressor with provable critical-point preservation (cpSZ) runs on the CPU at throughput far below the data-generation rates of modern GPU-based systems. We observe that, although preserving critical points is inherently a coupled and sequential constraint, it can be reformulated into independent parallel tasks, either on a per-block basis or, speculatively, on a per-point basis. We present FaCTz, the first GPU-based error-bounded lossy compressor that guarantees critical-point preservation. FaCTz provides a block-wise mode optimized for throughput and a speculative per-point mode optimized for compression ratio. Across three vector-field datasets, FaCTz preserves every critical point while achieving throughput of up to 60 GB/s, approximately two orders of magnitude (up to approximately 640x) faster than the multithreaded CPU implementation of cpSZ. Its speculative mode further improves the compression ratio by approximately a factor of two over the throughput-oriented mode.
Mingze Xia, Yuxiao Li, Sheng Di et al.· 0 citations
This first comprehensive study of data heterogeneity and lossless general-purpose compression performance for representative datasets from the PETRA III synchrotron radiation source provides a quantitative basis for future archival and storage decisions at PETRA III, its future successor, PETRA IV, and other large-scale scientific facilities.
M. Buschmann, Yannis Schumann, Christian Voss et al.· 0 citations
It is argued that replacing autoregressive LLMs with DLMs within the same compression framework could overcome the throughput bottleneck caused by their one-symbol-per-step limitation, and the newly proposed DLM-based framework advances the state of the art in lossless text compression.
Lossy compression is essential for managing massive scientific data, but per-element error bounds do not translate into bounds on downstream quantities of interest (QoIs) such as regional averages, neural network predictions, or multi-field derived quantities. We present TOPIQ, a statistical error-propagation framework that predicts QoI-level bias and uncertainty from compact compression metadata (less than 0.1% of original data). TOPIQ decomposes QoIs into primitive operators with closed-form propagation rules accounting for spatial error correlation and data-error coupling; new QoIs are supported by composition at runtime with no per-QoI derivation or retraining. Across 552 evaluations spanning 4 datasets, 3 compressors, 4 QoI families, and 8 error bounds, 93.1% of configurations achieve well-calibrated predictions. Pre-computed metadata enables post-hoc uncertainty quantification for arbitrary query regions at 56x-402x speedup over direct computation. A case study demonstrates integration into an AI-driven analysis pipeline with end-to-end confidence intervals for dynamically composed queries.
You-Yuan Liu, Bo Jiang, Taolue Yang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.