Skip to content
Preprint

Matryoshka Language Model Suites

Aug 2026 · 0 citations · 41 references
Computer Science

TL;DR

This paper improves both training and inference efficiency by stacking sub-models of increasing size into a single nested architecture trained end-to-end in a Matryoshka training framework that reduces the total parameter count, enables low-cost distillation from the largest to all smaller sub-models at every training step, and is well-suited for speculative decoding.

Abstract

Training a language model suite classically requires training each model separately and serving them independently. We improve both training and inference efficiency by stacking sub-models of increasing size into a single nested architecture trained end-to-end. This Matryoshka training framework reduces the total parameter count of the suite, enables low-cost distillation from the largest to all smaller sub-models at every training step, and is well-suited for speculative decoding as the draft model is contained within the verifier. We validate our approach by training a Matryoshka suite comprising 500M, 1.5B, and 3B sub-models. Our suite is on par with independently trained baselines on benchmark performance and validation and out-of-domain perplexities, while using 36% less training compute and improving the throughput of speculative decoding by 14-26%. We also ablate key architectural choices, offering guidance for building strong Matryoshka LM suites.

View source

Similar papers

Preprint Aug 2026

Projector Is All You Train

It is found that training only the projector is sufficient to achieve strong multimodal performance relative to existing baseline models and the authors' jointly trained MLLMs with the same encoder and backbone, and that joint training leads to undesirable drift in existing capabilities of the language model.

Nyx Iskandar, S. Selvan, Slater Victoroff · 0 citations
Preprint Aug 2026

Thinking at the Right Size: Amortized Distillation Across Post-Trained LLMs

ADAPT---Amortized Distillation Across Post-Trained LLMs is introduced, a framework for amortizing distillation across both axes of a model family: size and post-training variant, producing models for interpolated sizes across post-trained variants with a single distillation run.

Yan Zhou, Sara Kangaslahti, Jonathan Geuter et al. · 0 citations
Preprint Aug 2026

TokEval: A Tokenizer Evaluation Suite

TokEval is introduced, a framework of tokenizer evaluation metrics that goes beyond standard measures like fertility and compression rate to capture linguistically and structurally meaningful properties, e.g., UTF-8 character boundary integrity and digit place-value boundary alignment for mathematics.

Clara Meister · 1 citation
#machine learning Preprint Aug 2026

TokEval: A Tokenizer Evaluation Suite

Language model tokenizers are typically selected with minimal evaluation, despite the fact that their design choices directly impact model capabilities. This can be partly attributed to a limited understanding of which tokenizer properties affect which aspects of downstream performance. We introduce TokEval, a framewor...

Clara Meister · 0 citations
#artificial intelligence Preprint Sep 2026

Looped GPT-BERT: Trading Parameters for Computation in Small Language Modeling

When training data are limited, increasing parameter count is not the only way to improve language-model performance. A small parameter set, when repeatedly applied, can also deliver comparable performance. We study Looped GPT-BERT in the BabyLM 2026 Strict-small setting, combining GPT-BERT's masked next-token and caus...

Tingshuo Fan, Hong-Tao Mu, Tian-Yun Zhou et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.