In-context learning (ICL) enables a pretrained model to infer a task from demonstrations without updating its parameters. While much of the existing theory focuses on linear target functions, in this paper we study nonlinear cases by comparing two one-layer attention architectures on the same family of single-index tas...
Hao-Tian Gu, Yi-Zhou Xu, L. Zdeborová· 0 citations
Attention mechanisms are central to modern foundation models, yet their training dynamics remain poorly understood, especially when the attention matrices have extensive rank. In this work, we study attention-indexed models, a broad framework that can represent multi-layer and multi-head attention architectures. First,...
Yizhou Xu, M. Sagitova, Lenka Zdeborová et al.· 0 citations
A new error metric is introduced that precisely captures model vulnerability to consistent adversarial attacks -- perturbations that preserve the ground-truth labels, offering theoretical insight into the mechanisms underlying model sensitivity to adversarial attacks.
Matteo Vilucchio, Lenka Zdeborová, Bruno Loureiro· arXiv.org· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.