Preprint
Jul 2026
Data-Efficient Adaptation of LLMs via Attention Head Reweighting
Experiments on diverse open-source text classification datasets show that AHR can outperform standard baselines like LoRA when learning from limited samples, despite having 200-1000x fewer trainable parameters, as the authors' AHR only modifies ~0.0001% of the model's parameters.
Tuomas P. Oikarinen, Zixiao Chen, Charlotte Siska et al.
· 0 citations