This paper presents a dictionary architecture that integrates modules for morphological analysis, lemmatization, semantic clustering, and dictionary entry generation using large language models (LLMs) and proposes the first holistic conceptual architecture of an explanatory dictionary for Tajik that unifies classical lexicographic methods, language statistics, and generative capabilities of LLMs into a single system.
Abstract
This paper presents a conceptual framework for developing an electronic explanatory dictionary of the Tajik language using large language models (LLMs). The relevance of the work stems from the absence of a comprehensive digital lexicographic resource for Tajik that is comparable in functionality to dictionaries for high-resource languages, and from the limited adaptation of modern natural language processing technologies to low-resource language systems. Based on a systematic survey of existing linguistic, statistical, and corpus resources, we propose a dictionary architecture that integrates modules for morphological analysis, lemmatization, semantic clustering, and dictionary entry generation using LLMs. The choice of subword tokenization is justified by the agglutinative nature of Tajik morphology and its high morphological variability, along with a parameter-efficient fine-tuning (PEFT) strategy suitable for limited annotated data. The novelty of the work lies in proposing the first holistic conceptual architecture of an explanatory dictionary for Tajik that unifies classical lexicographic methods, language statistics, and generative capabilities of LLMs into a single system. The practical significance of the study is the formation of a methodological foundation for developing a full-featured electronic dictionary that can serve both as a lexicographic tool and as a core resource for machine translation, automatic summarization, sentiment analysis, and other applied NLP tasks. The paper is intended for specialists in computational linguistics, lexicography, and developers of natural language processing systems working with low-resource languages.
The section concludes with formal problem specification: given vocabulary V and corpus C, semantic analysis is formalized as a mapping problem preserving distributional properties, an optimization problem minimizing loss through gradient-based methods, and an evaluation problem assessing quality through semantic simila...
D. Akhmedjanova· Международный Журнал Теорети...· 0 citations
Evaluating language technology for low-resource languages faces a fundamental challenge: the scarcity of native benchmarks suitable for systematic assessment. For Faroese, no such evaluation frameworks exist. We address this gap by presenting the first benchmark suite for Faroese semantic understanding and grammatical...
Iben Nyholm Debess, Barbara Scalvini, Bolette S. Pedersen· International Conference on...· 0 citations
We present a systematic evaluation of Large Language Models (LLMs) in translating classical descriptions of phonological and morphophonological change from Old Indo-Aryan (Sanskrit) to Middle Indo-Aryan (MIA) into the standard notation of modern historical linguistics. Drawing on Vararuci’s Prākṛta Prakāśa (c. 4th ce...
V.S.D.S.Mahesh Akavarapu, Chinmay Dharurkar, Johannes Dellert et al.· Computational Linguistics· 0 citations
LLM-assisted sense assignment with a Serbian WordNet-based custom inventory, iterative inventory expansion, and expert validation is combined with a constrained JSON-formatted output to support the practical construction and refinement of sense-annotated resources in a low-resource setting.
Saša Petalinkar, R. Stanković, Milica Ikonić Nešić et al.· Intelligent Data Analysis· 0 citations
Inspicio is presented, an open-vocabulary retrieval pipeline that links tokens in context to synsets of the Open English WordNet (McCrae et al., 2020) without requiring any source-language inventory or mapping.
This study provides an initial basis for the development of Javanese semantic resources that are sensitive to speech-level variation and semantic complexity and has the potential to support future research in NLP applications in Javanese, including politeness identification, lexical normalization, word sense disambigua...
Musthofa Galih Pradana, Ridwan Raafi’udin, Nurul Afifah Arifuddin et al.· JOURNAL OF APPLIED INFORMATI...· 0 citations
Related blog posts
MIT News · Artificial Intelligence· news.mit.eduSep 24, 2026
A new method, called CW-Net, translates the reasoning process of an autonomous vehicle’s AI system into understandable concepts that explain its behavior.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.