Due to the complexity of neural network loss landscapes, optimization theory is forced to rely on idealized models, and there is generally a tradeoff between how theoretically tractable the model is, and how accurately it describes the true optimization dynamics. In this work, we stress test the simplest possible model...
Alexandru Meterez, Pranav Ajit Nair, Depen Morwani et al.· arXiv.org· 3 citations
State Space Models (SSMs) have emerged as a compelling alternative to Transformers, enabling sequence modeling with constant memory and linear compute. Although SSMs exhibit reasonable performance and favorable computational characteristics, they continue to lag behind Transformers on tasks that require in-context lear...
W. Tong, Aryo Lotfi, Emmanuel Abbe et al.· 0 citations
It is suggested that correlated prompts alter not only the effective sample size of in-context learning, but also which attention architectures are best matched to the task.
Mary I. Letey, Yue M. Lu, Cengiz Pehlevan et al.· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.