Root, Suffix, or Byte? A Reproducible Study of Sub Word Tokenization for a Morphologically Rich Language (Russian)
Tokenization is the first design choice in a language model and one of the most frequently overlooked for morphologically rich languages, as it directly shapes sequence length and how related word forms are segmented into reusable sub word units. Using Russian as a case study, we train sub word tokenizers from scratch...