Scaling Attention Head Analysis via Gradient-Based Attribution in Context-Aware Machine Translation
A gradient-based head attribution strategy where the Token-level Max-Margin loss is backpropagated to the attention maps enables a large-scale causal analysis of attention heads, making it suitable for LLMs.
Paweł Mąka, Yusuf Can Semerci, Jan Scholtes et al.
· 0 citations