Skip to content
Open access

Effects of Pivot Prompting and Text Type on LLM Translation Quality for the Low-Resource Chinese–Vietnamese Pair: Evidence from COMET and Human Evaluation

Jul 2026 · Applied Sciences · 0 citations · 34 references

TL;DR

The results show that the value of a prompting strategy is text-type-dependent and that fluent LLM output is not necessarily faithful, underscoring the need to pair automatic metrics with human evaluation when benchmarking low-resource translation.

Abstract

Large language models (LLMs) such as ChatGPT have advanced machine translation, but their quality on low-resource language pairs remains uneven and is typically assessed with automatic metrics alone. This study examines how prompting strategy and source-text type jointly affect LLM translation quality for the low-resource Chinese–Vietnamese pair and whether automatic and human assessments agree. Using a 2 × 3 mixed factorial design, we compared an English-mediated pivot strategy with a direct strategy across informative, expressive, and operative texts, evaluating 60 ChatGPT translations with COMET and with 15 bilingual readers who rated adequacy, fluency, faithfulness, trustworthiness, willingness to use, and need for revision. On COMET, pivoting significantly improved overall quality (p < 0.001), text type was the dominant factor (η2 = 0.91; informative > operative > expressive), and strategy interacted with text type, with the largest pivot gain for expressive texts. Human ratings reproduced this ordering but diverged sharply for expressive texts: although COMET favoured the pivot output, readers reported that pivoting raised fluency yet substantially reduced faithfulness (4.10 → 1.77 on a 7-point scale) and were unwilling to accept it. These results show that the value of a prompting strategy is text-type-dependent and that fluent LLM output is not necessarily faithful, underscoring the need to pair automatic metrics with human evaluation when benchmarking low-resource translation.

Read PDF

Similar papers

Can We Triage LLM Translation Errors in Classical Texts Without Human References? Source Novelty, GEMBA Scoring, and Budgeted Review through Pali-to-English Translation

A budgeted workflow is proposed, combining source novelty, peer disagreement, and stronger candidate-aware judging to allocate human review through Pali-to-English translation, indicating that evaluator strength matters beyond the prompt alone.

Máté Metzger · 1 citation
Review Open access Aug 2026

Accuracy and error patterns of ChatGPT-4o for real-time English-Nepali voice translation: A cross-sectional field evaluation in rural Nepal

ChatGPT-4o demonstrated potential for real-time English-Nepali communication but also produced errors that could alter interpretation of participant responses, which support cautious use for low-stakes conversational exchange and human verification when errors could affect research validity, clinical decisions, or part...

A. Mandich, S. Koirala, S. Westen et al. · 0 citations
Review Open access Aug 2026

Evaluating Cultural Balance and Representation in ChatGPT and DeepSeek-Generated EFL Materials

The findings show that cultural inclusion does not ensure balanced representation and the proposed review criteria help teachers and curriculum developers evaluate variation, specificity, stereotyping risk, and classroom suitability before using AI-generated materials.

Idrees, Yong-Zhi Liu · 0 citations
Preprint Aug 2026

Investigating the Influence of Prompt and Response Languages on LLM Content Generation

This study examines how prompt and response language influence the behavior of large language models. Using five models, we evaluated answers to 68 non translation questions across four language conditions: English to English, English to Norwegian, Norwegian to Norwegian, and Norwegian to English. After removing refuse...

Thi Thanh Huyen Nguyen, Mai Khoi Tieu, M. Riegler et al. · 0 citations
Open access Jul 2026

Large Language Models for Japanese–Croatian Translation: Human Evaluation and Macroeconomic Implications

Large Language Models (LLMs) are increasingly used for translation, yet their value depends on preserving meaning rather than producing fluent output. This study evaluates seven LLMs on Japanese–Croatian translation, a low-resource, typologically distant language pair. Using rubric-based human evaluation of adequacy, f...

Ratomir Karlović, Mieta Bobanović Dasko, Irena Srdanović · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.