Skip to content
Preprint

Grammar Engineering Meets LLMs: Development of Cantonese and Irish ParGram Treebanks

Aug 2026 · 1 citation · ⚡ 1 influential · 60 references
Computer Science

TL;DR

The development of Cantonese and Irish treebanks within the Parallel Grammar (ParGram) Project is presented, where linguistic parallelism is maintained at an abstract functional level and the methodological potential and limitations of using multilingual LLMs to support grammar engineering are investigated.

Abstract

Grammar engineering requires expertise in linguistic formalism and computational implementation, especially in parallel grammar projects that balance cross-linguistic consistency with language-specific properties. This paper presents the development of Cantonese and Irish treebanks within the Parallel Grammar (ParGram) Project, where linguistic parallelism is maintained at an abstract functional level. We also investigate the methodological potential and limitations of using multilingual LLMs to support grammar engineering, focusing on Cantonese-Irish translation and the generation of formal syntactic structures using OpenAI's gpt-oss-120b model. The results show that translation performance was generally unsatisfactory and unaffected by prompt language. For syntactic structure generation, the model produced some structurally meaningful outputs, but performed poorly on tasks requiring cross-linguistic abstraction. Nonetheless, LLM-generated outputs may still offer some reference value by suggesting alternative analyses and (partially) capturing predicate-argument relations. Overall, our findings highlight both the potential and limitations of using LLMs in collaborative grammar engineering, while underscoring the continued importance of expert-driven analysis and verification.

View source

Similar papers

Preprint Aug 2026

How Useful are LLMs for Grammar Engineering? Cantonese ParGram Resources and Controlled Experimental Evaluation with English Baselines

The results characterize both the capabilities and limitations of current LLMs for potential integration into AI-assisted expert workflows: LLMs may support intermediate stages of grammar development, but human linguistic expertise remains central to analysis, validation, and refinement.

Chit-Fung Lam · 0 citations
#artificial intelligence Preprint Sep 2026

Multilingual in Name Only? Cultural and Linguistic Weaknesses of LLMs in Urdu

Multilingual large language models (LLMs) are increasingly used for open-ended text generation, yet their behaviour in low-resource languages remains poorly understood. In this work, we question how correct and reliable is the generation of multilingual LLMs when used for the task of story generation. We consider Urdu...

Farah Adeeba, A. Khan, Rajesh Bhatt et al. · 0 citations
Open access Aug 2026

Using Large Language Models in Formalizing Classical Linguistic Descriptions: A Case Study in Middle Indo-Aryan Sound Change

We present a systematic evaluation of Large Language Models (LLMs) in translating classical descriptions of phonological and morphophonological change from Old Indo-Aryan (Sanskrit) to Middle Indo-Aryan (MIA) into the standard notation of modern historical linguistics. Drawing on Vararuci’s Prākṛta Prakāśa (c. 4th ce...

V.S.D.S.Mahesh Akavarapu, Chinmay Dharurkar, Johannes Dellert et al. · 0 citations
Review Open access 2025

Large Language Models Under Evaluation: An Acceptability, Complexity And Coherence Assessment In Italian

The results suggest that, although fine-tuned transformers outperform all GPT models, GPT-4 represents a significant improvement over third-generation GPT models and poses a challenge to the poverty of the stimulus hypothesis.

Cristiano Chesi, F. Vespignani, Roberto Zamparelli · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.