Preprint
Jul 2026
Complexity-Guided Component-wise Initialization for Language Model Pretraining
It is suggested that pretrained spectra are useful diagnostics of trained model structure, but that effective reuse likely requires preserving richer information than component-wise scale and singular-value shape, while coarse spectral matching alone is not a reliable optimization strategy.
Konstantin Garbers, Nicholas Oh
· 0 citations