KITE: KV-Invariant Transformer Expansion for Efficient Agentic LLM Scaling
Scaling a language model is not only a question of final quality: the architectural choice determines how much computation is spent during training, prompt processing, and autoregressive decoding to achieve certain model quality. An ideal model architecture should lower all above computation costs to facilitate scaling...