Skip to content

A Layered Taxonomy for Chinese Learner Grammatical Error Annotation

Sep 2026 · 0 citations · 26 references
Computer Science

TL;DR

A layered scheme linking computational Chinese grammatical error correction (CGEC) with pedagogical error analysis and a preliminary consistency study in which five large language models apply it to a sample support the layered approach while identifying category boundaries requiring further refinement.

Abstract

Grammatical error annotation in Chinese learner writing requires labels that are both consistent and linguistically meaningful. This paper proposes a layered scheme linking computational Chinese grammatical error correction (CGEC) with pedagogical error analysis. The scheme first identifies character- and punctuation-level orthographic errors, labeling them by edit operation and subtype. Other errors receive a three-layer core label combining edit operation, linguistic domain, and part of speech, with optional Chinese-specific extensions for aspect, modality, comparison, argument structure, and complements. Drawing on CGEC resources, learner-error taxonomies, and Mandarin grammar, the taxonomy is evaluated through a coverage analysis of automatically extracted MuCGEC edits and a preliminary consistency study in which five large language models apply it to a sample. The results support the layered approach while identifying category boundaries requiring further refinement.

View source

Similar papers

Open access Aug 2026

BERT-based Automatic Error Correction System for Chinese Learners

Improved robustness and explainability for automatic Chinese learner error correction is demonstrated by an alignment consistency loss to ensure character-level consistency, and the combination of word-segmentation augmentation and multi-reference soft-label training to reduce conflicts caused by segmentation differenc...

Y.-L. Diao, W. Gao · 0 citations
Open access Sep 2026

Automatic grammar error correction model in English writing teaching

This paper developed a automatic grammar error correction (GEC) model based on the sequence-to-sequence (seq2seq) model. It adopted a dual-encoding structure composing of a syntactic encoder and a semantic encoder, and introduced a hybrid attention mechanism in the decoder. Finally, the results were output through the...

Yan Liu · 0 citations
Open access Sep 2026

Word and Insertion-Gap Error Detection in English Learner Writing: A Cross-Corpus Study of Pretrained and Syntactic Features

Grammatical error detection must identify both problematic source words and locations where material is missing. This study evaluates these decisions separately and examines whether explicit dependency features improve classifiers built on frozen pretrained representations. An audited W&I+LOCNESS partition supplies 31,...

Xiao Yang · 0 citations
Open access 2026

PARSEME 2.0 Multilingual Corpus of Multiword Expressions

We present edition 2.0 of the PARSEME multilingual corpus annotated for multiword expressions (MWEs), resulting from efforts of the PARSEME community towards universality-driven modeling of idiomaticity. With respect to previous editions, we extend the annotation scope to all syntactic MWE categories: verbal, nominal,...

Agata Savary, Manon Scholivet, Carlos Ramisch et al. · 1 citation
Open access Sep 2026

Assessing mechanical, morphological, and syntactic error patterns in written English grammatical accuracy among English language learners at SIMAD University, Somalia

This study examined mechanical, morphological, and syntactic error patterns in the written English grammatical accuracy of first-year learners at SIMAD University. Grounded in Error Analysis theory, this study used a descriptive quantitative design and analyzed 223 in-class written essays (100–150 words) to identify, c...

Mustaf Ahmed Wehlie, Ahmed Abdullahi Mohamud · 0 citations
Open access Sep 2026

An Investigation into the Omission of Chinese Object Errors by Japanese and Korean Learners — An Analysis Based on the HSK Dynamic Writing Corpus

Using the HSK Dynamic Composition Corpus 3.0, this study examines 541 instances of object omission in Chinese compositions written by Japanese and Korean learners. The screened sample contains 277 instances from Japanese learners and 264 from Korean learners. Nominal objects account for 487 cases (90.02%), including 32...

Wen-Pan Guo · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.