A layered scheme linking computational Chinese grammatical error correction (CGEC) with pedagogical error analysis and a preliminary consistency study in which five large language models apply it to a sample support the layered approach while identifying category boundaries requiring further refinement.
Abstract
Grammatical error annotation in Chinese learner writing requires labels that are both consistent and linguistically meaningful. This paper proposes a layered scheme linking computational Chinese grammatical error correction (CGEC) with pedagogical error analysis. The scheme first identifies character- and punctuation-level orthographic errors, labeling them by edit operation and subtype. Other errors receive a three-layer core label combining edit operation, linguistic domain, and part of speech, with optional Chinese-specific extensions for aspect, modality, comparison, argument structure, and complements. Drawing on CGEC resources, learner-error taxonomies, and Mandarin grammar, the taxonomy is evaluated through a coverage analysis of automatically extracted MuCGEC edits and a preliminary consistency study in which five large language models apply it to a sample. The results support the layered approach while identifying category boundaries requiring further refinement.
Improved robustness and explainability for automatic Chinese learner error correction is demonstrated by an alignment consistency loss to ensure character-level consistency, and the combination of word-segmentation augmentation and multi-reference soft-label training to reduce conflicts caused by segmentation differenc...
Y.-L. Diao, W. Gao· Advanced Electromagnetics· 0 citations
This paper developed a automatic grammar error correction (GEC) model based on the sequence-to-sequence (seq2seq) model. It adopted a dual-encoding structure composing of a syntactic encoder and a semantic encoder, and introduced a hybrid attention mechanism in the decoder. Finally, the results were output through the...
Yan Liu· Discover Artificial Intellig...· 0 citations
Grammatical error detection must identify both problematic source words and locations where material is missing. This study evaluates these decisions separately and examines whether explicit dependency features improve classifiers built on frozen pretrained representations. An audited W&I+LOCNESS partition supplies 31,...
Xiao Yang· Archives des sciences: a mul...· 0 citations
We present edition 2.0 of the PARSEME multilingual corpus annotated for multiword expressions (MWEs), resulting from efforts of the PARSEME community towards universality-driven modeling of idiomaticity. With respect to previous editions, we extend the annotation scope to all syntactic MWE categories: verbal, nominal,...
Agata Savary, Manon Scholivet, Carlos Ramisch et al.· International Conference on...· 1 citation
This study examined mechanical, morphological, and syntactic error patterns in the written English grammatical accuracy of first-year learners at SIMAD University. Grounded in Error Analysis theory, this study used a descriptive quantitative design and analyzed 223 in-class written essays (100–150 words) to identify, c...
Mustaf Ahmed Wehlie, Ahmed Abdullahi Mohamud· Asian-Pacific Journal of Sec...· 0 citations
Using the HSK Dynamic Composition Corpus 3.0, this study examines 541 instances of object omission in Chinese compositions written by Japanese and Korean learners. The screened sample contains 277 instances from Japanese learners and 264 from Korean learners. Nominal objects account for 487 cases (90.02%), including 32...
Wen-Pan Guo· Global Education Ecology· 0 citations
Related blog posts
MIT News · Artificial Intelligence· news.mit.eduSep 24, 2026
A new method, called CW-Net, translates the reasoning process of an autonomous vehicle’s AI system into understandable concepts that explain its behavior.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.