MDM-Prime-v2: Binary Encoding and Index Shuffling Enable Scaling of Diffusion Language Models
This work analyzes the optimal design of the subtokenizer that minimizes MDM-Prime training objective and develops MDM-Prime-v2, a masked diffusion language model which incorporates Binary Encoding and Index Shuffling.
Chen-Hao Chao, Wei-Fang Sun, Jun-Wei Quan et al.
· 0 citations