Skip to content

Removable and Irreducible: A Token-Cost Ledger for the Multilingual Tokenization Tax

Jul 2026 · 0 citations
Computer Science

TL;DR

A token-cost ledger is assembled that splits each language's cost, at fixed parallel content, into a removable coding redundancy, a residual coding slack, an intrinsic-content term, and an orthogonal, irreducible grapheme-to-phoneme term that governs the multimodal rather than the text cost.

Abstract

Large language models pay a well-documented tax on non-English text: the same content costs several times more tokens, and because attention is quadratic in sequence length, far more compute. We ask how much of this tax is removable. Framing the token layer as source coding -- transformer compute is monotone in sequence length, whose per-atom floor is the Shannon rate $H/\log_2 V$, an object already applied to tokenizers in prior work -- we assemble a token-cost ledger that splits each language's cost, at fixed parallel content, into a removable coding redundancy, a residual coding slack, an intrinsic-content term, and an orthogonal, irreducible grapheme-to-phoneme term that governs the multimodal rather than the text cost. On FLORES-200 across eight languages, a production tokenizer costs up to $8.9\times$ more tokens for Indic scripts than for English; a script-matched code trained on $1,012$ sentences removes a median $64\%$ of that excess (bootstrap 95\% CI $[0.638, 0.647]$), and a script-fair information floor shows the intrinsic content differs by under $6\%$ -- the tax is representational, not informational. A constructed code removes $98\%$ of a controlled source's redundancy, and the token tax implies up to $79\times$ attention cost. We are explicit about scope and failure: this is compute-and-memory accounting, not a model-quality claim; we neither measure nor claim the cross-lingual direction of the orthographic term; and our matched code is a conservative small-data demonstration. We contribute the unifying ledger, the removable-versus-intrinsic attribution, and an open one-command harness.

View source

Similar papers

#natural language process... Preprint Oct 2026

Counting and Min-Cost Encoding for Tokenization in Large Language Models

Mainstream large language models rely on a tokenizer to encode text into a token sequence. Different tokenizers may yield token sequences of substantially different lengths for the same text. With a fixed model architecture, shorter token sequences correspond to lower inference time. We propose a tokenizer training app...

Shu-Ming Shi, Xiang Zhang, Hao-Yi Yu et al. · 0 citations
#natural language process... Preprint Oct 2026

More Than Words: Compositional Tokenization for Efficient Language Models

Language models process and generate text sequentially in token units, and the tokenizer determines how much text each inference step covers. Under standard tokenization, a short English phrase such as"On the table."is usually produced as four separate predictions for the preposition (On), article (the), noun (table),...

Yuval Reif, Guy Kaplan, Roy Schwartz · 0 citations
#natural language process... Preprint Sep 2026

Fewer Words, Not Fewer Tokens: Measuring the Sanskrit Tokenization Penalty per Proposition

Sanskrit fuses case, number, person and tense into word endings and chains clauses into compounds, so it is information-dense per word. Whether that density survives subword tokenization is a separate question, to be asked per unit of meaning rather than per word. On identical FLORES-200 devtest content, Sanskrit costs...

Dev Sharma · 0 citations
#natural language process... Preprint Oct 2026

Byte Language Models: Scaling, Emergent Abstractions, and Information Allocation

Tokenizer-free language models remove the inductive bias of fixed tokenizers by modeling text directly as bytes, but the resulting longer sequences substantially increase computation and eliminate explicit text abstractions. We ask whether this additional computation can be useful, and whether standard Transformers can...

Jie Wang, Shi-Wei Luo, Qi Zhang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

The Price of Token Boundaries: Compression Certificates and Prediction

Pre-tokenisation restricts which text fragments can become prediction units, but its compression cost is obscured when tokenisers are compared only under the same boundaries. We measure this cost by bounding the minimum token count from both sides, with and without a regular-expression boundary rule. Nonnegative prices...

Yu-Hao Du, Shu-Nian Chen · 0 citations
#artificial intelligence Preprint Sep 2026

Breaking the Token Ceiling: Distilling Smaller, Stronger Byte Models

Small models are made more capable through distillation from a larger one that shares their tokenization scheme. However, do distilled byte and token models behave similarly in terms of scaling trends as compute and data increases? To enable this comparison, we introduce two variants to efficiently convert token logits...

Kalyani Marathe, Artidoro Pagnoni, Tomasz Limisiewicz et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.