It is found that for a significant fraction of inputs, the LLM's distribution agrees with the ENTD almost perfectly, and the agreement generally increases with model scale and training compute, but there is a long tail of input sequences where the LLM and ENTD differ significantly.
Abstract
In this paper, we study the connection between an LLM's output distribution and the data used to train it. Specifically, we study the degree to which an LLM's next-token distribution agrees with the empirical next-token distribution (ENTD) given the context in the training data. The ENTD is an appealing target because it is the unrestricted global minimizer of the next-token cross entropy loss used for pretraining, as well as an easily interpretable function of the pretraining corpus. We find that for a significant fraction of inputs, the LLM's distribution agrees with the ENTD almost perfectly, and the agreement generally increases with model scale and training compute. Nevertheless, there is a long tail of input sequences where the LLM and ENTD differ significantly, and we examine several possible sources of this discrepancy across the transformer architecture, training procedure, and finite-sample noise in the ENTD estimate itself. More broadly, we hope our findings will encourage more work on ``data-centric mechanistic interpretability,''a complement to standard mechanistic interpretability that opens the black box of how model behaviors arise from the data, rather than how they are encoded in the learned weights.
The empirical results indicate that the reported gains depend heavily on the experiment setup and the choice of random seed, although the trend of stable gains is confirmed with 128-Dyck pretraining of small models with the Llama tokenizer for most of the examined languages.
Sofiia Riazhskykh, Nam Luu, Ondrej Bojar· 0 citations
The degradation rate across neural models, both sentence embeddings and decoder-only LLMs, is studied, and how consistent it is depends on the scale of the noise: under word-level noise, models with very different architectures decline along nearly the same curve, while under character-level noise they separate.
The experimental results show that even 130M-parameter models benefit from including the MTP task in the pre-training objective, and hold even under severe data constraints, as demonstrated on both zero-shot benchmarks and downstream tasks.
The concavity of the matrix-based conditional entropy functional is proved, which makes the resulting entropy-constrained projection a convex optimization problem, and a scalable mirror-descent algorithm for its implementation is developed.
This work identifies a curse of ambiguity: in large language models, and more broadly in all neural networks that produce discrete probability distributions, the more ambiguous a next-token distribution is, the harder it is to learn accurately.
Nicolas Zucchet, Hyun Dong Lee, Scott W. Linderman· 2 citations
This work finds that repetition counts tuned on smaller proxy models with the same \(\mathrm{TPP}\) can provide a practical estimate for larger models, and suggests that repetition counts tuned on smaller proxy models with the same \(\mathrm{TPP}\) can provide a practical estimate for larger models.
Jingwei Li, Xinran Gu, Rui Dai et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.