Improved robustness and explainability for automatic Chinese learner error correction is demonstrated by an alignment consistency loss to ensure character-level consistency, and the combination of word-segmentation augmentation and multi-reference soft-label training to reduce conflicts caused by segmentation differences.
Abstract
Current automatic error correction methods for Chinese learners often focus on superficial word- or sentence-level processing and are affected by inconsistent annotation standards, resulting in limited generalization to real learner texts and causing misalignment or overcorrection. To improve grammatical correctness, semantic fidelity, and instructional relevance, this paper constructs a multi-granularity lexical representation method. Using characters as the basic unit, the method integrates three types of information: characters, words, and pinyin. These representations are learned through projection processing, concatenated into a unified embedding, and fed into a pre-trained Chinese BERT encoder. A dual-task head consisting of sequence labeling and lightweight generation is deployed on the shared BERT encoder for collaborative optimization. Key innovations include an alignment consistency loss to ensure character-level consistency, and the combination of word-segmentation augmentation and multi-reference soft-label training to reduce conflicts caused by segmentation differences. Experimental results show that the method achieves granular alignment consistency above 0.800, corrected-original sentence similarity of 0.888, and an overcorrection rate as low as 0.041. These results demonstrate improved robustness and explainability for automatic Chinese learner error correction.
A layered scheme linking computational Chinese grammatical error correction (CGEC) with pedagogical error analysis and a preliminary consistency study in which five large language models apply it to a sample support the layered approach while identifying category boundaries requiring further refinement.
This paper addresses the issues of unstable recognition, semantic misalignment, and excessive inference costs that often occur in "OCR + machine translation" under complex backgrounds, vertical-horizontal mixed layouts, and low-resolution conditions by constructing a set of automated translation optimization framework...
Sha Wang· International Conference on...· 0 citations
This paper proposes an alignment-aware multi-granularity tagging framework. First, this method uses a cross-lingual pre-trained model to encode source and target language contexts jointly while explicitly modeling-level bilingual correspondences via a learnable soft alignment layer. Second, a gated local enhancement mo...
Bi-Juan Wang, Lingli Zhu, Hongli Wen· International Journal of Inf...· 0 citations
For a long time, there has been a lack of effective cultural semantic conversion and automatic detection methods for term inconsistency in the English Chinese translation of cultural heritage terminology. This paper constructs an automatic error detection model for cultural heritage terminology translation based on Bid...
This paper addresses the task of automatic identification of German word forms and constructs an end-to-end model framework based on deep neural networks that employs character-level and subword-level dual-channel feature representations, and combines encoder-decoder architecture, scaled dot-product attention, and posi...
Bo Wang· International Conference on...· 0 citations
The research offers a Lotus Effect-Attention-based Bi-directional Gated Recurrent Unit (LE-Att-Bi-GRU) deep learning model for automatic translation quality assessment that improves semantic representation by incorporating a lotus-inspired division method that decreases noise and focuses essential semantic cues.