Skip to content
Preprint

How Useful are LLMs for Grammar Engineering? Cantonese ParGram Resources and Controlled Experimental Evaluation with English Baselines

Aug 2026 · 0 citations · 37 references
Computer Science

TL;DR

The results characterize both the capabilities and limitations of current LLMs for potential integration into AI-assisted expert workflows: LLMs may support intermediate stages of grammar development, but human linguistic expertise remains central to analysis, validation, and refinement.

Abstract

This paper presents new Cantonese ParGram resources and evaluates LLMs for knowledge-driven grammar engineering within a controlled experimental paradigm. Using Cantonese ParGram resources as gold standards, with corresponding English baselines, we investigate whether OpenAI's gpt-oss-120b and GPT-5.4 can generate machine-processable grammars from sentences and target formal structures under systematically varied prompting conditions. GPT-5.4 outperformed gpt-oss-120b, while grammars generated from target formal structures generally outperformed those generated from sentences. Although both models could generate locally plausible phrase-structure rules, lexical entries, and templates, they often struggled to coordinate interacting formal constraints, especially in multi-construction settings. The results characterize both the capabilities and limitations of current LLMs for potential integration into AI-assisted expert workflows: LLMs may support intermediate stages of grammar development, but human linguistic expertise remains central to analysis, validation, and refinement. The study also contributes new Cantonese symbolic grammatical resources.

View source

Similar papers

Preprint Aug 2026

Grammar Engineering Meets LLMs: Development of Cantonese and Irish ParGram Treebanks

The development of Cantonese and Irish treebanks within the Parallel Grammar (ParGram) Project is presented, where linguistic parallelism is maintained at an abstract functional level and the methodological potential and limitations of using multilingual LLMs to support grammar engineering are investigated.

Chit-Fung Lam, Elaine Uí Dhonnchadha · 1 citation · ⚡1

A Factorial Study of Synthetic Data Generation for Low-Resource Machine Translation using Grammar Books

A pipeline that uses large language models to extract grammatical rules, example sentences, and lexicons from grammar books and generate synthetic parallel corpora for fine-tuning-rather than feeding grammar content into prompts at inference time, as in prior work is introduced.

V. Ravikumar, Sina Ahmadi, L. Jäger et al. · 0 citations
Open access Aug 2026

Using Large Language Models in Formalizing Classical Linguistic Descriptions: A Case Study in Middle Indo-Aryan Sound Change

We present a systematic evaluation of Large Language Models (LLMs) in translating classical descriptions of phonological and morphophonological change from Old Indo-Aryan (Sanskrit) to Middle Indo-Aryan (MIA) into the standard notation of modern historical linguistics. Drawing on Vararuci’s Prākṛta Prakāśa (c. 4th ce...

V.S.D.S.Mahesh Akavarapu, Chinmay Dharurkar, Johannes Dellert et al. · 0 citations
#artificial intelligence Preprint Sep 2026

What Do We Expect from LLMs? Mapping the Design of LLM Benchmarks

This work systematically map 14,767 papers introducing or updating evaluation resources from arXiv submissions between January 2022 and August 2026, using staged screening and automated full-text coding to examine changes in target systems and domains, evaluation materials and conditions, and scoring mechanisms.

Chao Wang · 0 citations
Open access Aug 2026

Optimizing Context and Cost in LLM‐Based Unit Test Generation: A Study on External Dependency Retrieval Strategies

A systematic empirical study of multiple strategies for context enrichment and optimization in LLM‐based unit test generation, conducted on seven diverse projects (three open‐source and four proprietary industrial systems), encompassing 261 distinct methods establish this optimized context strategy as a cost‐effective...

Javier Ferrer, Francisco Chicano · 0 citations
Review Open access 2025

Large Language Models Under Evaluation: An Acceptability, Complexity And Coherence Assessment In Italian

The results suggest that, although fine-tuned transformers outperform all GPT models, GPT-4 represents a significant improvement over third-generation GPT models and poses a challenge to the poverty of the stimulus hypothesis.

Cristiano Chesi, F. Vespignani, Roberto Zamparelli · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.