Context-Sensitive N-Gram Word Partitioning for Improving the Quality of Turkish Word Embeddings
The proposed approach provides a language-agnostic, context-sensitive segmentation mechanism that can complement language processing methods such as lemmatization, morphological analysis, and stemming and indicate task-dependent and generally limited improvements over traditional token-based word-embedding extraction.