Skip to content

Author

Junming Zhang

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Jul 2026

Memory Decoder at Scale: A Pretrained, Parametric Long-Term Memory

This work presents Memory Decoder at Scale, scaling memory models up to 6.9B parameters and pretraining them on 300B tokens, demonstrating that independently scaling pretrained memory offers a more parameter efficient path to improving language model performance.

Rubin Wei, Jiaqi Cao, Jiarui Wang et al. · 1 citation