Preprint
Jul 2026
Tokenizing Crosslingual Homographs
This work proposes a simple tokenizer-level intervention based on language cues: language-specific characters replacing initial characters of shared-vocabulary words, reducing common identity during vocabulary construction, and suggests that adding lightweight language information at the tokenizer level is a promising direction for further exploration.
Rotem Brillant, Yuval Pinter
· 0 citations