Skip to content

In-Place Tokenizer Expansion for Pre-trained LLMs

Jul 2026 · arXiv.org · Vol abs/2607.15232 · 0 citations · 52 references
Computer Science

Abstract

A tokenizer fixed at the start of pre-training allocates vocabulary in proportion to the pre-training corpus, reflecting the deployment priorities at that time. When those priorities shift, languages added later are split into many more tokens per word, which can raise latency, compute, and energy consumption for users of those languages. Cloud models can afford a broad vocabulary because the embedding and LM-head matrices are a small fraction of their parameters. On a compact model those matrices are a material share of per-token decode bandwidth, so on-device models ship small vocabularies and accept fragmentation outside a fixed language set. We present tokenizer expansion, an in-place recipe for upgrading a pre-trained model's tokenizer when the model producer controls its design. We continue the existing tokenizer's BPE merges on a multilingual corpus, so most source tokens carry over unchanged as single tokens and every new token has an exact decomposition into source tokens. We copy the carried-over embedding rows unchanged and initialize new rows as the mean of their source sub-token embeddings. A two-stage adaptation, embedding-only training then full-model continued pre-training, recovers source-checkpoint quality. We apply the recipe to a continued pre-trained checkpoint of LFM2-8B-A1B, an 8B-parameter Mixture-of-Experts model, to help produce LFM2.5-8B-A1B with a 128K tokenizer. The expanded tokenizer encodes Hindi and Vietnamese in roughly $2.4\times$ and $2.6\times$ fewer tokens than the source (up to $4.0\times$ on Thai). Combining these reductions with the measured per-token cost of the larger vocabulary, we estimate a $2.2$-$3.7\times$ per-character decode speedup for these languages across our reference devices. We release the model weights and the expanded tokenizer, and report the negative findings that shaped the recipe.

View source

Similar papers

#artificial intelligence Preprint Aug 2026

Nested Byte-Level Vocabularies Are Cheap to Deploy and Expensive to Share: A Pre-Registered Negative Result

Slicing a byte-level BPE tokenizer allows one language model to operate at several vocabulary sizes, use a control token to indicate the active size, and be deployed at any trained size by slicing its embedding and output head, yielding a falsifiable prediction for future work.

Christos Koutsiaris · 0 citations
Conference Open access Sep 2026

Subword Tokenization for Low- and Medium-Resource Languages: A Systematic Evaluation

Subword tokenization is a standard technique for pre-trained language models, mapping text into sequences of tokens from a fixed-size vocabulary. Despite its widespread use, the impact of tokenization algorithms and vocabulary sizes on downstream performance remains underexplored, particularly for low- and medium-resou...

J. Daðason, H. Loftsson · 0 citations
#natural language process... Preprint Aug 2026

When Tokenizers Fail: Byte-Level Chunking for Zero-Shot Transfer to Low-Resource Languages

This paper initializes byte embeddings directly from the subword representations of a frozen base model, applies a chunk alignment loss to project dynamically grouped byte chunks toward precomputed subword targets, and interleave lightweight part-of-speech supervision to guide boundary detection.

Sanjeev Kumar, Atsuki Yamaguchi, Nikolaos Aletras · 0 citations
#natural language process... Preprint Jul 2026

Removable and Irreducible: A Token-Cost Ledger for the Multilingual Tokenization Tax

A token-cost ledger is assembled that splits each language's cost, at fixed parallel content, into a removable coding redundancy, a residual coding slack, an intrinsic-content term, and an orthogonal, irreducible grapheme-to-phoneme term that governs the multimodal rather than the text cost.

Madhulatha Mandarapu, Sandeep Kunkunuru · 0 citations
Preprint Aug 2026

Instability of LLM Pre-Pretraining: It Doesn't Always Help. An Investigation on Multiple Languages

The empirical results indicate that the reported gains depend heavily on the experiment setup and the choice of random seed, although the trend of stable gains is confirmed with 128-Dyck pretraining of small models with the Llama tokenizer for most of the examined languages.

Sofiia Riazhskykh, Nam Luu, Ondrej Bojar · 0 citations
#artificial intelligence Preprint Aug 2026

The Curse of Multilinguality in Lexical Normalization

This work tests whether a language's typological distance from the others predicts its ideal number of co-training languages, and finds no dependable rule: any apparent relationship rests on a couple of languages and does not hold up.

Saman Rahbar · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.