Skip to content
Preprint

Tied Trit-Planes: Constraining PTQTP to a Uniform Nine-Level Quantizer, with a Persistent Folded Format for Disk-Streamed Mixture-of-Experts Serving

Aug 2026 · 0 citations · 32 references
Computer Science

TL;DR

This work is the first to impose that identity as a constraint inside PTQTP's solver, a known balanced-ternary identity, and applies it to the routed experts of DeepSeek-V4-Flash-0731, a 284B-A13B mixture-of-experts model, quantizing in one shot from the released MXFP4 expert weights and streaming experts from SSD on a 64 GB laptop.

Abstract

PTQTP decomposes LLM weight matrices into two ternary (trit) planes with two free per-group scales. Tying the scales to a fixed ratio of three collapses the decomposition into a single uniform nine-level quantizer, a known balanced-ternary identity. To our knowledge, at the time of writing, this work is the first to impose that identity as a constraint inside PTQTP's solver. The two trit planes then fold losslessly into one 4-bit code plane that we make the persistent serving representation: disk bytes, expert-cache bytes, and kernel input are the same 4.0625-bits/weight blocks, consumed in one integer dot pass. For this conjunction (ratio-3 nine-level code, CPU-SIMD kernels, SSD expert streaming, identical persistent bytes) we likewise found no precedent. We apply this to the routed experts of DeepSeek-V4-Flash-0731, a 284B-A13B mixture-of-experts model, quantizing in one shot from the released MXFP4 expert weights and streaming experts from SSD on a 64 GB laptop. Against a 4.5-bit Q4_K baseline, measured one process per fixture with an expert-lossless anchor arm as reference control, the tied model matches the official serving API on 5/5 fixtures at step 0 (Q4_K: 4/5) and 12/14 captured continuation steps (11/14), scores 86 vs. 84 on a 100-item MMLU subset, decodes 6.7% faster in decode phase, and ships 9% smaller files: no detected fidelity difference at these small evaluation sizes, and every fixture-level difference between the arms traces to a single measured near-tie cell. The tied fit nevertheless shows higher weight-reconstruction error and worse perplexity, a measured dissociation between proxy metrics and reference fidelity. A cumulative trunk-ternarization ladder and bitwise-pinned aarch64/x86-64 kernels complete the report. All code, formats, and evaluation artifacts are open source in the fucina inference stack.

View source

Similar papers

Review Aug 2026

Transforms for LLM Quantization: The Great Inversion and Format Co-Design

This work identifies and formalizes the principle that organizes the Great Inversion, the Great Inversion: allocation-flexible coding rewards energy concentration, whereas the grouped shared-scale quantization a deployed matrix instruction performs rewards within-group flattening.

Ehsan Jokar · 0 citations
Preprint Aug 2026

SchurQuant: Groupwise Discrete Optimization for Layer-Wise LLM Quantization

SCHUROPT is introduced, which analytically eliminates the suffix's optimal continuous response, yielding an exact groupwise quadratic with Schur-complement curvature, and achieves the highest mean zero-shot accuracy among the evaluated backpropagation free PTQ baselines.

Gunjun Lee, Sehwan Son, Younjoo Lee et al. · 0 citations
#machine learning Preprint Sep 2026

JARQ: Joint Alternating Refinement for Quantization

Group-wise post-training quantizers for large language models round weights onto a grid that is not refit to the resulting integer codes. We show that this leaves accuracy on the table: the best grid depends on the codes, input correlations couple the errors of different groups, and useful code changes often involve ma...

Xin-Yu Wang, Sicheng Lyu, Xiao-Wen Chang · 0 citations
#artificial intelligence Preprint Sep 2026

Per-Matrix Optimality Is Not Enough: Three-Level Optimization for Low-Rank LLM Compression

A three-level chain that widens optimization scope from individual matrices to Transformer blocks to the full model, with improvements only in perplexity and compression fidelity; downstream accuracy remains well below the dense model.

Hui-Cheng Zhang, Xi-Yao Feng, Ze-Tong Li et al. · 0 citations
#artificial intelligence Preprint Oct 2026

Cascadia: Resident 975B MoE Inference on Eleven AI PCs

Mixture-of-experts models make nearly trillion-parameter capacity accessible with sparse per-token computation, provided that the serving system can distribute the weights and coordinate their execution. We present Cascadia's resident execution of Inkling, a 975B-total/41B-active-parameter model, on eleven Intel Core U...

Tate Berenbaum, Matias Parij, M. Venkatachalam · 0 citations
Open access Aug 2026

Coding-Level Evaluation of a Kronecker-Sequence Interleaver Under Synthetic Three-Dimensional Correlated Fault Models

This study evaluates a fixed Kronecker-sequence interleaver under controlled synthetic three-dimensional correlated-fault models. Spatial-cluster, column-correlated, and bit-plane-dependent probability fields are used as coding-level abstractions and are not calibrated device measurements. The K-IPA mapping is compared...

Qiulin He, Dongliang Zhang, Ru Lu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.