Squeezing More from Limited Data with Recursive Transformers
Recursion Transformers are studied, reusing a shared block across depth to scale compute, together with factorized embeddings to reduce vocabulary-map parameters, and find that they outperform standard Transformers at 10M and 100M words, while remaining competitive with BabyLM Challenge 2025 winners.