Scaling Attention Head Analysis via Gradient-Based Attribution in Context-Aware Machine Translation
A gradient-based head attribution strategy where the Token-level Max-Margin loss is backpropagated to the attention maps enables a large-scale causal analysis of attention heads, making it suitable for LLMs.