Preprint
Aug 2026
LegoLM: Structured Weight Sharing for Large Language Models
It is discovered that outlier dominance grows with model scale: full replacement at K=128 degrades GPT-2 small but catastrophically degrades Mistral-7B by +1,134,279%, while selective replacement at p=99% rescues both models to under +15%.
Joseph Bingham
· 0 citations