LZ Penalty: An information-theoretic repetition penalty for autoregressive language models
The LZ penalty is introduced, a penalty specialized for reducing degenerate repetitions in autoregressive language models without loss of capability and without instances of degenerate repetition, and enables state-of-the-art open-source reasoning models to operate with greedy decoding without loss of capability and without instances of degenerate repetition.
Antonio A. Ginart, Naveen Kodali, Jason Lee et al.
· 0 citations