Language models are typically pretrained from random initialization. Recent work challenges this convention, showing that a brief warm-up on abstract, algorithmically generated data can provide a better starting point for subsequent learning of natural language. In this paper, we show that in small language models, suc...
Zachary Shinnick, Hemanth Saratchandran, Damien Teney et al.· 0 citations
This paper proposes conditioned initialization, a principled scheme that initializes attention weights to improve the spectral properties of the attention layer and shows that conditioned initialization can potentially reduce the condition number of the attention Jacobian, leading to more stable optimization.
The results show that standard transformers are rarely a local optimum in the space of architectures, and suggests that there may be room for improved architectures that better support multiple capabilities simultaneously, such as fluency and robust reasoning.