Skip to content

Can We Triage LLM Translation Errors in Classical Texts Without Human References? Source Novelty, GEMBA Scoring, and Budgeted Review through Pali-to-English Translation

Sep 2026 · 1 citation · 47 references
Computer Science

TL;DR

A budgeted workflow is proposed, combining source novelty, peer disagreement, and stronger candidate-aware judging to allocate human review through Pali-to-English translation, indicating that evaluator strength matters beyond the prompt alone.

Abstract

As large language models become capable translators of classical texts, a key challenge is deciding which outputs need expert review when no human reference exists. This study tests reference-free error triage through Pali-to-English translation. Three LLMs translated 15,493 passages. Five signals were compared: source novelty, source-candidate embedding distance, peer-translation disagreement, English-to-Pali backtranslation, and no-reference GEMBA scoring. Signals were calibrated on a 3,000-item reference-informed LLM-adjudicated sample and checked against a 500-item author-adjudicated anchor. Human references supported calibration and validation only; they were never used to compute the risk signals. Source novelty was a useful source-side risk prior but not a per-candidate error detector. Peer disagreement and backtranslation provided secondary signal. The strongest method was no-reference GEMBA scoring by a panel of models generally regarded as stronger than the translators: reviewing the top 10% by GEMBA risk captured 81.6% of panel-major errors in the calibration set. GEMBA also remained the best reference-free signal against the author anchor. A same-tier panel, with self-scoring excluded, remained useful but performed worse, indicating that evaluator strength matters beyond the prompt alone. A budgeted workflow is proposed, combining source novelty, peer disagreement, and stronger candidate-aware judging to allocate human review. Transfer to other classical languages, including Latin, Ancient Greek, and Sanskrit, remains to be tested.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

TermJudge: A Document-Level Metric Judging, Not Counting, Terminology in Machine Translation Evaluation

Existing automatic metrics for evaluating terminological use in machine translation (MT) penalise any divergence from a fixed reference, conflating translation errors with the valid terminological variation that human translators routinely produce. We introduce TermJudge, a document-level terminology metric that assign...

Nicolas M. Dahan, Franccois Yvon, Rachel Bawden · 0 citations
#natural language process... Preprint Sep 2026

Discourse Dependency: A Continuous Criterion for Translation Difficulty

This work argues that one meaningful and currently unmeasured axis is referential reach, the distance a segment must look back into its document to resolve the entities and pronouns it contains, and formalizes this as discourse dependency (DDP), a metric-free, source-side measure computed from named entity re-mentions...

Ahrii Kim, Chanjun Park, Seong-heum Kim · 0 citations

Evaluating Multilingual Sentence Embeddings for Translation Error Detection:An English--Greek Contrastive Study

The results show complementary error-sensitivity profiles: multilingual sentence embeddings provide useful semantic adequacy signals but are better suited as components of broader translation-evaluation frameworks than as standalone metrics.

Eleftherios Kalogeros, Athanasios Ntalakas, M. Gergatsoulis et al. · 0 citations
#natural language process... Preprint Sep 2026

In the Blind: Building Pseudo-References for MT Evaluation

The WMT26 General MT task evaluates systems on 10 language pairs that have no human references (neither translated from scratch nor post-edited from MT output by humans) and how the pseudo-references were built is described.

Diptesh Kanojia, Chi-Kiu Lo, Archchana Sindhujan et al. · 0 citations

English Translation Quality Estimation And Fine-Grained Error Pattern Recognition Based On Multi-Source Corpus Augmentation And Large Language Models

Translation quality estimation (QE) must support both sentence-level scoring and local error diagnosis under limited, heterogeneous supervision. This study proposes MCA-LLM-QEER, which combines leakage-controlled multi-source corpus augmentation, source-language/crosslingual/ target-language semantic evidence, explicit...

Wei Wang, Wei-Feng Liu · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.