This paper improves both training and inference efficiency by stacking sub-models of increasing size into a single nested architecture trained end-to-end in a Matryoshka training framework that reduces the total parameter count, enables low-cost distillation from the largest to all smaller sub-models at every training step, and is well-suited for speculative decoding.
Abstract
Training a language model suite classically requires training each model separately and serving them independently. We improve both training and inference efficiency by stacking sub-models of increasing size into a single nested architecture trained end-to-end. This Matryoshka training framework reduces the total parameter count of the suite, enables low-cost distillation from the largest to all smaller sub-models at every training step, and is well-suited for speculative decoding as the draft model is contained within the verifier. We validate our approach by training a Matryoshka suite comprising 500M, 1.5B, and 3B sub-models. Our suite is on par with independently trained baselines on benchmark performance and validation and out-of-domain perplexities, while using 36% less training compute and improving the throughput of speculative decoding by 14-26%. We also ablate key architectural choices, offering guidance for building strong Matryoshka LM suites.
It is found that training only the projector is sufficient to achieve strong multimodal performance relative to existing baseline models and the authors' jointly trained MLLMs with the same encoder and backbone, and that joint training leads to undesirable drift in existing capabilities of the language model.
Nyx Iskandar, S. Selvan, Slater Victoroff· 0 citations
ADAPT---Amortized Distillation Across Post-Trained LLMs is introduced, a framework for amortizing distillation across both axes of a model family: size and post-training variant, producing models for interpolated sizes across post-trained variants with a single distillation run.
Yan Zhou, Sara Kangaslahti, Jonathan Geuter et al.· 0 citations
TokEval is introduced, a framework of tokenizer evaluation metrics that goes beyond standard measures like fertility and compression rate to capture linguistically and structurally meaningful properties, e.g., UTF-8 character boundary integrity and digit place-value boundary alignment for mathematics.
The concavity of the matrix-based conditional entropy functional is proved, which makes the resulting entropy-constrained projection a convex optimization problem, and a scalable mirror-descent algorithm for its implementation is developed.
Language model tokenizers are typically selected with minimal evaluation, despite the fact that their design choices directly impact model capabilities. This can be partly attributed to a limited understanding of which tokenizer properties affect which aspects of downstream performance. We introduce TokEval, a framewor...
When training data are limited, increasing parameter count is not the only way to improve language-model performance. A small parameter set, when repeatedly applied, can also deliver comparable performance. We study Looped GPT-BERT in the BabyLM 2026 Strict-small setting, combining GPT-BERT's masked next-token and caus...
Tingshuo Fan, Hong-Tao Mu, Tian-Yun Zhou et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.