Effects of Pivot Prompting and Text Type on LLM Translation Quality for the Low-Resource Chinese–Vietnamese Pair: Evidence from COMET and Human Evaluation
The results show that the value of a prompting strategy is text-type-dependent and that fluent LLM output is not necessarily faithful, underscoring the need to pair automatic metrics with human evaluation when benchmarking low-resource translation.
Abstract
Large language models (LLMs) such as ChatGPT have advanced machine translation, but their quality on low-resource language pairs remains uneven and is typically assessed with automatic metrics alone. This study examines how prompting strategy and source-text type jointly affect LLM translation quality for the low-resource Chinese–Vietnamese pair and whether automatic and human assessments agree. Using a 2 × 3 mixed factorial design, we compared an English-mediated pivot strategy with a direct strategy across informative, expressive, and operative texts, evaluating 60 ChatGPT translations with COMET and with 15 bilingual readers who rated adequacy, fluency, faithfulness, trustworthiness, willingness to use, and need for revision. On COMET, pivoting significantly improved overall quality (p < 0.001), text type was the dominant factor (η2 = 0.91; informative > operative > expressive), and strategy interacted with text type, with the largest pivot gain for expressive texts. Human ratings reproduced this ordering but diverged sharply for expressive texts: although COMET favoured the pivot output, readers reported that pivoting raised fluency yet substantially reduced faithfulness (4.10 → 1.77 on a 7-point scale) and were unwilling to accept it. These results show that the value of a prompting strategy is text-type-dependent and that fluent LLM output is not necessarily faithful, underscoring the need to pair automatic metrics with human evaluation when benchmarking low-resource translation.
A budgeted workflow is proposed, combining source novelty, peer disagreement, and stronger candidate-aware judging to allocate human review through Pali-to-English translation, indicating that evaluator strength matters beyond the prompt alone.
ChatGPT-4o demonstrated potential for real-time English-Nepali communication but also produced errors that could alter interpretation of participant responses, which support cautious use for low-stakes conversational exchange and human verification when errors could affect research validity, clinical decisions, or part...
A. Mandich, S. Koirala, S. Westen et al.· medRxiv· 0 citations
The findings show that cultural inclusion does not ensure balanced representation and the proposed review criteria help teachers and curriculum developers evaluate variation, specificity, stereotyping risk, and classroom suitability before using AI-generated materials.
Idrees, Yong-Zhi Liu· International "Journal of Ac...· 0 citations
This study examines how prompt and response language influence the behavior of large language models. Using five models, we evaluated answers to 68 non translation questions across four language conditions: English to English, English to Norwegian, Norwegian to Norwegian, and Norwegian to English. After removing refuse...
Thi Thanh Huyen Nguyen, Mai Khoi Tieu, M. Riegler et al.· 0 citations
Large Language Models (LLMs) are increasingly used for translation, yet their value depends on preserving meaning rather than producing fluent output. This study evaluates seven LLMs on Japanese–Croatian translation, a low-resource, typologically distant language pair. Using rubric-based human evaluation of adequacy, f...