Preprint
Aug 2026
Divergent large language model predictions from convergent representations in ambiguous word pairs
This work investigates how decoder-only transformers resolve lexical ambiguity through layer-by-layer analysis of three models spanning three parameter sizes, finding that representations become maximally distinct in middle layers, then partially reconverge in late layers, while the KL divergence between their next-token predictions reaches its maximum in the final layers.
K. Scott, Narun Pat, Veronica Liesaputra
· 0 citations