Skip to content

Author

Mengdi Wang

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Jul 2026

DeepLoop: Depth Scaling for Looped Transformers

The results show that stable recurrent depth requires residual scaling rules that account for parameter visits, not only nominal layer count, and DeepLoop is neutral when no physical block is revisited and improves validation loss and downstream accuracy once recurrent depth is activated.

Shu-zhen Li, Yi-Fan Zhang, Jiacheng Guo et al. · 2 citations
#machine learning Preprint Aug 2026

Fast Weight Attention for Continual Learning

This framework separates temporal alignment, plasticity, forgetting, and bounded rehearsal in recurrent sequence models, together with numerically stable positive-decay renormalization, to remain competitive in language modeling and improve length extrapolation on variable-digit addition.

Yi-Fan Zhang, Steve Ta, Jasper Zhang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.