Jul 2026
Multi-Head Attention Residuals
Multi-Head Attention Residuals (MHAR) is introduced: the routing query is reshaped into H per-subspace heads, each with its own softmax over the depth history, and a direct probe of the trained queries confirms that learned subspace disagreement is the underlying driver.
Cheng Luo, Zefan Cai, Jun-Jie Hu
· arXiv.org · 1 citation