A structured and comprehensive review of Sino-Tibetan languages PoS taggers
Abstract
Parts of Speech (PoS) tagging is a fundamental activity in Natural Language Processing (NLP). It consists in assigning an appropriate grammatical category to each word in a text, such as noun, verb, adjective or adverb. It serves as a crucial preprocessing step in many NLP applications such as machine translation, information retrieval, text summarization and sentiment analysis. PoS tagging has been significantly improved, however, obtaining high accuracy for different languages is still an active research challenge. Researchers have developed different ways throughout the years starting from simple lexical methods to rule-based and statistical techniques and most recently hybrid, deep learning and transformer-based models which provide better performance. This study reviews the current PoS tagging approaches comprehensively and discusses the various strategies, its advantages and limitations. Special emphasis is made on languages of the Sino-Tibetan family, many of which are considered as low-resource, due to the lack of annotated corpora and linguistic resources. As shown by the survey, relatively little work has been done on PoS tagging for these languages compared to the far more extensively studied languages. To fill this gap, this study examines existing studies on PoS tagging of several Sino-Tibetan languages, explains the datasets and models used, and compares their stated performance. Furthermore, the article discusses the limits of the prior approaches and highlights the major research gaps and future research objectives to create more robust machine learning, deep learning, hybrid and transformer-based PoS tagging systems for low-resource Sino-Tibetan languages.