Preprint
Jul 2026
Memory Decoder at Scale: A Pretrained, Parametric Long-Term Memory
This work presents Memory Decoder at Scale, scaling memory models up to 6.9B parameters and pretraining them on 300B tokens, demonstrating that independently scaling pretrained memory offers a more parameter efficient path to improving language model performance.
Rubin Wei, Jiaqi Cao, Jiarui Wang et al.
· 1 citation