Skip to content

Author

Hemanth Saratchandran

We have 3 of 66 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#machine learning Preprint Sep 2026

Lasting Effects of Abstract Pretraining Beyond Perplexity

Language models are typically pretrained from random initialization. Recent work challenges this convention, showing that a brief warm-up on abstract, algorithmically generated data can provide a better starting point for subsequent learning of natural language. In this paper, we show that in small language models, suc...

Zachary Shinnick, Hemanth Saratchandran, Damien Teney et al. · 0 citations
#machine learning Preprint Sep 2026

Conditioned Initialization for Attention

This paper proposes conditioned initialization, a principled scheme that initializes attention weights to improve the spectral properties of the attention layer and shows that conditioned initialization can potentially reduce the condition number of the attention Jacobian, leading to more stable optimization.

Hemanth Saratchandran, Simon Lucey · 0 citations
Jul 2026

Can Transformers Really Do It All? On the Compatibility of Inductive Biases Across Tasks

The results show that standard transformers are rarely a local optimum in the space of architectures, and suggests that there may be room for improved architectures that better support multiple capabilities simultaneously, such as fluency and robust reasoning.

Damien Teney, Liangze Jiang, Hemanth Saratchandran et al. · 2 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.