Aug 2026· 1 citation· ⚡ 1 influential· 60 references
Computer Science
TL;DR
The development of Cantonese and Irish treebanks within the Parallel Grammar (ParGram) Project is presented, where linguistic parallelism is maintained at an abstract functional level and the methodological potential and limitations of using multilingual LLMs to support grammar engineering are investigated.
Abstract
Grammar engineering requires expertise in linguistic formalism and computational implementation, especially in parallel grammar projects that balance cross-linguistic consistency with language-specific properties. This paper presents the development of Cantonese and Irish treebanks within the Parallel Grammar (ParGram) Project, where linguistic parallelism is maintained at an abstract functional level. We also investigate the methodological potential and limitations of using multilingual LLMs to support grammar engineering, focusing on Cantonese-Irish translation and the generation of formal syntactic structures using OpenAI's gpt-oss-120b model. The results show that translation performance was generally unsatisfactory and unaffected by prompt language. For syntactic structure generation, the model produced some structurally meaningful outputs, but performed poorly on tasks requiring cross-linguistic abstraction. Nonetheless, LLM-generated outputs may still offer some reference value by suggesting alternative analyses and (partially) capturing predicate-argument relations. Overall, our findings highlight both the potential and limitations of using LLMs in collaborative grammar engineering, while underscoring the continued importance of expert-driven analysis and verification.
The results characterize both the capabilities and limitations of current LLMs for potential integration into AI-assisted expert workflows: LLMs may support intermediate stages of grammar development, but human linguistic expertise remains central to analysis, validation, and refinement.
Multilingual large language models (LLMs) are increasingly used for open-ended text generation, yet their behaviour in low-resource languages remains poorly understood. In this work, we question how correct and reliable is the generation of multilingual LLMs when used for the task of story generation. We consider Urdu...
Farah Adeeba, A. Khan, Rajesh Bhatt et al.· 0 citations
We present a systematic evaluation of Large Language Models (LLMs) in translating classical descriptions of phonological and morphophonological change from Old Indo-Aryan (Sanskrit) to Middle Indo-Aryan (MIA) into the standard notation of modern historical linguistics. Drawing on Vararuci’s Prākṛta Prakāśa (c. 4th ce...
V.S.D.S.Mahesh Akavarapu, Chinmay Dharurkar, Johannes Dellert et al.· Computational Linguistics· 0 citations
The results suggest that, although fine-tuned transformers outperform all GPT models, GPT-4 represents a significant improvement over third-generation GPT models and poses a challenge to the poverty of the stimulus hypothesis.
Cristiano Chesi, F. Vespignani, Roberto Zamparelli· Italian Journal of Computati...· 1 citation
It is shown that translation is even more modular than previously assumed and that the output language production in translation processes is actually further separable into a syntax and a surface language process.
M. Sonkin, Tanja Baeumel, Daniil Gurgurov et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.