Training nGPT
A practical training recipe for the normalized Transformer and its evaluation on modern hybrid Mamba-2--Transformer Mixture-of-Experts models shows that the 30B-total-parameter nGPT model reaches the same validation loss using approximately half as many training tokens.
I. Loshchilov, Boris Ginsburg
· 0 citations