Preprint
Jul 2026
Wiola 13M, a Gated Spiral Attention Architecture for Parameter Efficient Small Language Models
Wiola is presented, a decoder only language model whose novelty is concentrated in three drop in components of every layer, and it is proved that the gated attention admits an exact and numerically verified equivalence between full sequence training and cached autoregressive decoding, so that no approximation is introduced at inference time.
A. Chowdhury, Praveen Oosa, V. Reddy
· 0 citations