Jul 2026· Digital Scholarship in the Humanities· 0 citations· 20 references
TL;DR
This study highlights the potential of NLP tools to streamline the semiautomatic annotation process, reducing the reliance on extensive linguistic expertise and manual effort, and paving the way for broader applications in digital humanities research.
Abstract
Building syntactically annotated corpora, such as treebanks, for historical languages is a challenging yet vital task in digital humanities, as it underpins linguistic analysis and facilitates a range of interdisciplinary research. However, the scarcity of annotated data and the need for extensive expertise in historical linguistics make this process particularly demanding. In this study, we explore the potential of cross-lingual natural language processing (NLP) techniques as a semiautomatic solution for treebank construction in low-resource historical languages. We use Middle High German (MHG) as a case study. Leveraging the linguistic continuity and structural similarities between MHG and Modern German (MG), we effectively utilize the extensive MG treebank resources to develop a constituency parsing system tailored for MHG. Specifically, to design a semiautomatic system that integrates automatic annotation with manual validation, we explore two cross-lingual transfer techniques: zero-shot transfer and delexicalization; the latter removes lexical information to focus on syntactic structure. In our experiments, we first train parsers on MG treebanks, and then transfer them to MHG using the two cross-lingual transfer techniques. The delexicalization method achieves a parsing performance of 67.3 per cent in terms of F1-score. This performance significantly surpasses the zero-shot cross-lingual method by a margin of 28.6 percentage points. These investigations validate the effectiveness and feasibility of cross-lingual transfer techniques for historical language treebank construction. This study highlights the potential of NLP tools to streamline the semiautomatic annotation process, reducing the reliance on extensive linguistic expertise and manual effort, and paving the way for broader applications in digital humanities research.
LLM-assisted sense assignment with a Serbian WordNet-based custom inventory, iterative inventory expansion, and expert validation is combined with a constrained JSON-formatted output to support the practical construction and refinement of sense-annotated resources in a low-resource setting.
Saša Petalinkar, R. Stanković, Milica Ikonić Nešić et al.· Intelligent Data Analysis· 0 citations
Large language models (LLMs) are increasingly used to classify, label, summarize, and interpret large text collections, creating new possibilities for corpus linguistics. Their capacity for zero-shot and few-shot instruction following could reduce the cost of linguistic annotation and extend analysis beyond the categories handled by conventional part-of-speech taggers, parsers, and dictionary-based tools. At the same time, LLM outputs are probabilistic, prompt-sensitive, model-dependent, and potentially biased, raising fundamental questions about measurement validity, annotation reliability, and reproducibility. This article critically synthesizes foundational corpus-annotation principles with recent evidence on LLM-based text annotation and develops a validated human–LLM workflow for corpus research. The framework distinguishes token-, span-, sentence-, document-, and discourse-level annotation; requires a human-coded gold sample before large-scale deployment; treats prompt design as part of the annotation manual; and evaluates accuracy, precision, recall, F1, inter-annotator agreement, stability across repeated runs, subgroup performance, and error types. Recent studies show that LLMs can approach or exceed crowd-worker performance on some well-specified classification tasks, but that performance varies substantially across datasets, languages, models, prompts, text lengths, and annotation complexity. Span-level annotation and context-dependent semantic or pragmatic coding remain particularly challenging. The article therefore argues against unvalidated full automation and proposes selective automation, disagreement-based human adjudication, model/version documentation, and preservation of raw outputs. For corpus linguistics, the strongest near-term use of LLMs is as flexible annotators within a transparent, theory-driven, and auditable pipeline rather than as replacements for linguistic expertise. The resulting framework supports scalable corpus annotation while preserving the empirical principles on which corpus-based linguistic inference depends.
Maria Ibrar· International Journal Of Lit...· 0 citations
A pipeline that uses large language models to extract grammatical rules, example sentences, and lexicons from grammar books and generate synthetic parallel corpora for fine-tuning-rather than feeding grammar content into prompts at inference time, as in prior work is introduced.
V. Ravikumar, Sina Ahmadi, L. Jäger et al.· arXiv.org· 0 citations
This approach addresses the data scarcity problem in Indonesian SRL by leveraging the availability of annotated English-language corpora by leveraging the availability of annotated English-language corpora.
Bariza Haqi, M. L. Khodra· Journal of ICT Research and...· 0 citations
We present a systematic evaluation of Large Language Models (LLMs) in translating classical descriptions of phonological and morphophonological change from Old Indo-Aryan (Sanskrit) to Middle Indo-Aryan (MIA) into the standard notation of modern historical linguistics. Drawing on Vararuci’s Prākṛta Prakāśa (c. 4th century CE) and English translation of Bhāamaha’s commentary (c. 6th century CE), we construct a dataset of 470 phonological and morphophonological rules extracted from these sources using LLMs followed by manual curation. Out of these rules, we compile a benchmark of 216 sound-change rules with example reflexes. Our pipeline integrates Optical Character Recognition (OCR) of non-digitized historical texts, manual gold-standard curation, and LLM-based translation of rule descriptions into contemporary phonological rule notation. Evaluation across several recent LLMs shows accuracies up to 86% on this challenging formalization task. Models incorporating explicit reasoning mechanisms consistently outperform non-reasoning variants, underscoring the importance of reasoning in linguistic formalization. Error analyses reveal systematic weaknesses in modeling complex conditioning environments. We further show that the extracted sound laws generalize well across a broader range of MIA languages. Overall, this work illustrates how contemporary LLMs can engage with millennia-old linguistic scholarship by systematically translating and structuring classical rule descriptions into modern formal representations.
V.S.D.S.Mahesh Akavarapu, Chinmay Dharurkar, Johannes Dellert et al.· Computational Linguistics· 0 citations
We present edition 2.0 of the PARSEME multilingual corpus annotated for multiword expressions (MWEs), resulting from efforts of the PARSEME community towards universality-driven modeling of idiomaticity. With respect to previous editions, we extend the annotation scope to all syntactic MWE categories: verbal, nominal, adjectival, adverbial and functional. We cover 17 languages, of which 7 are new. The annotation process is based on cross-lingually unified guidelines, phrased as decision diagrams over linguistic tests, and a typology of 18 MWE categories. The corpus contains almost 5 million tokens, over 250,000 sentences and 140,000 MWE annotations. The applicability of the corpus is tested in baseline experiments with a prompt-based MWE identification system. Results show that generic large language models do not encode sufficient knowledge to solve the MWE identification task.
Agata Savary, Manon Scholivet, Carlos Ramisch et al.· International Conference on...· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.