A Tibetan language model (TLM) integrates grammatical relationships and morphological verb features that outperforms existing methods on ASR tasks and proposes a suffix feature awareness method to model functional word continuation relationships in sentences.
Abstract
Language models (LMs) are essential for automatic speech recognition (ASR). Existing research on LMs has focused primarily on major world languages such as English and Chinese, with findings difficult to adapt to other languages. Tibetan, as the official language of Tibet Autonomous Region of China, has unique key properties—grammar and morphological verbs play crucial roles in understanding—that cannot be effectively modelled by existing LMs. We propose a Tibetan language model (TLM) integrates grammatical relationships and morphological verb features. Our main contributions include: (i) analysing the importance of Tibetan grammar and morphological verbs and their impact on LMs. (ii) Proposing a suffix feature awareness method to model functional word continuation relationships in sentences. (iii) Proposing an adaptive weighting method to address rare morphological verbs, which are predominantly low-frequency words in Tibetan. Experiments show that the method considering grammatical relationships achieves 4.8% lower perplexity (PPL) than state-of-the-art method. The morphological verb weighting method achieves 7.4% lower PPL, and the method combining grammar and morphological verbs achieves 9.8% lower PPL. Our approach also outperforms existing methods on ASR tasks.
This study examines the four major word classes (viz. noun, verb, adjective and adverb) in Lohorung, a Kirat Rai (Tibeto-Burman) language of north-eastern Nepal, based on their semantic, syntactic, and morphological properties, based on Givón’s formal and functional approach. Data were collected through elicitation wit...
Diwas Rai· Journal of Indigenous Knowle...· 0 citations
Parts of Speech (PoS) tagging is a fundamental activity in Natural Language Processing (NLP). It consists in assigning an appropriate grammatical category to each word in a text, such as noun, verb, adjective or adverb. It serves as a crucial preprocessing step in many NLP applications such as machine translation, info...
Bedawati Basumatary, Shikhar Kumar Sarma, Kuwali Talukdar et al.· Frontiers in Artificial Inte...· 0 citations
Prefixation is a morphological process which studies how morphemes are affixed in the beginning of their host words for either grammatical functions or for lexical expansion and derivation. This paper investigates aspects of the grammar of prefixation in English and Igbo Languages. It adopts a comparative theoretical f...
Okolo N Patrick, Florence Nne Agwu· World Journal of Advanced Re...· 0 citations
This paper analyses various recent state-of-the-art variants of large language models (LLMs) and neural machine translation (NMT) for Indian languages in comparison to statistical machine translation (SMT) and tackles key questions, such as idiomatic expressions, morphologically complex grammar or the scarceness of par...
Jayanand A. Kamble, Shivajirao M. Jadhav, V. J. Kadam· International Journal of Inf...· 0 citations
Part-of-speech (POS) tagging in Arabic is hard because its rich root-and-pattern morphology and the absence of short vowels make one unvoweled string compatible with many categories. This paper presents Jawhar, a hybrid framework that couples a high-performance morphological analyser with contextual reranking using a p...
Mohamed Bouzahir, A. A. Abdelouahad, M. Nabil· Information· 0 citations
Lexical bundles and lexical phrases (p-frames) are multi-word sequences that have been widely researched. However,
function word frames (FWFs) — p-frames with fixed function words and variable content words — remain underexplored in
argumentative writing. Thai students struggle with function words because their f...
Ali Wang, Ridwan Wahid· Australian Review of Applied...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.