Preprint
Aug 2026
Mixture of Training: Recombining Small-Scale Scaffolded Pretraining Runs into a Larger Language Model
Mixture of Training is introduced, a scaffolded modular pre-training procedure that partitions a target Transformer into contiguous layer blocks, trains each block inside a frozen pretrained aligner scaffold, and then recomposes the trained blocks with an optional short end-to-end adaptation pass to study whether scaffolded sub-runs can act as reusable training units.
Mohammed Sabry, Sean Augenstein, Keith Rush et al.
· 0 citations