Skip to content
Preprint

Scaling Domain Data Repetition in LLM Pretraining

Aug 2026 · 1 citation · 40 references
Computer Science

TL;DR

This work finds that repetition counts tuned on smaller proxy models with the same \(\mathrm{TPP}\) can provide a practical estimate for larger models, and suggests that repetition counts tuned on smaller proxy models with the same \(\mathrm{TPP}\) can provide a practical estimate for larger models.

Abstract

As large language models scale, their training-token budgets must also increase to maintain an appropriate tokens-per-parameter ratio (\(\mathrm{TPP}\)). However, high-quality domain data is much harder to scale than general web data. As model size and the training-token budget increase, its fraction in the training mixture tends to decrease. Repeating the available high-quality data provides an effective way to counteract this dilution, but excessive repetition may lead to overfitting. We study this trade-off under practical LLM scaling, where the training-token budget grows proportionally with model size. For a fixed domain, we first find that, surprisingly at a fixed \(\mathrm{TPP}\), the optimal repetition count mildly increases with model size. Across different domains, we find that the optimal repetition count is strongly negatively correlated with the final validation loss of a domain: domains with lower loss can generally benefit from more repetitions. In contrast, the amount of unique domain data is only weakly related to the optimal repetition count. These findings suggest that repetition counts tuned on smaller proxy models with the same \(\mathrm{TPP}\) can provide a practical estimate for larger models.

View source

Similar papers

Preprint Aug 2026

Squeezing More from Limited Data with Recursive Transformers

Recursion Transformers are studied, reusing a shared block across depth to scale compute, together with factorized embeddings to reduce vocabulary-map parameters, and find that they outperform standard Transformers at 10M and 100M words, while remaining competitive with BabyLM Challenge 2025 winners.

Serdar Gülbahar, Lukas Edman, Alexander Fraser · 0 citations

The Pitfalls of Maximum Likelihood Training of TPMs for Controllable Language Generation

This work finds that likelihood-trained TPMs can result in failed generations due to overly large corrections to the LM’s logits, and trains TPMs with LM-aligned objectives that better align with the LM token-probability space.

Hanzhang Liu, William Zhao, Zi-Lei Shao et al. · 0 citations
Jul 2026

Convolution for Large Language Models

These results support depthwise convolution as a lightweight complement to self-attention for modeling short-range token interactions and suggest that the convolution makes repeated token IDs more sensitive to their immediate context.

Yu-Chuan Tian, Yingte Shu, Wei He et al. · 0 citations
#machine learning Preprint Sep 2026

Data Scarcity and Model Sparsity: Mixtures-of-Experts Overfit More to Repeated Data

As the supply of human-written text is exhausted, it has become standard practice to repeat language model training data. Prior work has studied data repetition for densely activated Transformers, but the effects of data repetition remains largely unexplored for recently dominant sparse architectures such as Mixture-of...

Atindra Jha, Margaret Li, J. Leskovec et al. · 1 citation
Preprint Aug 2026

Why Large Language Models Fail at Tabular Prediction

The results show that the LLM's capability dissolves with dimension in a way no noise-corrupted classical learner mimics - which explains why LLMs, so capable elsewhere, keep losing to fifty-year-old baselines on tables, while leaving the mechanism of the prediction as an open question.

M. Garnelo, Wojciech M. Czarnecki · 1 citation

Babies Learn to Look Ahead: Multi-Token Prediction in Small LMs

The experimental results show that even 130M-parameter models benefit from including the MTP task in the pre-training objective, and hold even under severe data constraints, as demonstrated on both zero-shot benchmarks and downstream tasks.

Ansar Aynetdinov, Alan Akbik · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.