Can We Triage LLM Translation Errors in Classical Texts Without Human References? Source Novelty, GEMBA Scoring, and Budgeted Review through Pali-to-English Translation
A budgeted workflow is proposed, combining source novelty, peer disagreement, and stronger candidate-aware judging to allocate human review through Pali-to-English translation, indicating that evaluator strength matters beyond the prompt alone.
Abstract
As large language models become capable translators of classical texts, a key challenge is deciding which outputs need expert review when no human reference exists. This study tests reference-free error triage through Pali-to-English translation. Three LLMs translated 15,493 passages. Five signals were compared: source novelty, source-candidate embedding distance, peer-translation disagreement, English-to-Pali backtranslation, and no-reference GEMBA scoring. Signals were calibrated on a 3,000-item reference-informed LLM-adjudicated sample and checked against a 500-item author-adjudicated anchor. Human references supported calibration and validation only; they were never used to compute the risk signals. Source novelty was a useful source-side risk prior but not a per-candidate error detector. Peer disagreement and backtranslation provided secondary signal. The strongest method was no-reference GEMBA scoring by a panel of models generally regarded as stronger than the translators: reviewing the top 10% by GEMBA risk captured 81.6% of panel-major errors in the calibration set. GEMBA also remained the best reference-free signal against the author anchor. A same-tier panel, with self-scoring excluded, remained useful but performed worse, indicating that evaluator strength matters beyond the prompt alone. A budgeted workflow is proposed, combining source novelty, peer disagreement, and stronger candidate-aware judging to allocate human review. Transfer to other classical languages, including Latin, Ancient Greek, and Sanskrit, remains to be tested.
Existing automatic metrics for evaluating terminological use in machine translation (MT) penalise any divergence from a fixed reference, conflating translation errors with the valid terminological variation that human translators routinely produce. We introduce TermJudge, a document-level terminology metric that assign...
Nicolas M. Dahan, Franccois Yvon, Rachel Bawden· 0 citations
This work argues that one meaningful and currently unmeasured axis is referential reach, the distance a segment must look back into its document to resolve the entities and pronouns it contains, and formalizes this as discourse dependency (DDP), a metric-free, source-side measure computed from named entity re-mentions...
Ahrii Kim, Chanjun Park, Seong-heum Kim· 0 citations
The results show complementary error-sensitivity profiles: multilingual sentence embeddings provide useful semantic adequacy signals but are better suited as components of broader translation-evaluation frameworks than as standalone metrics.
Eleftherios Kalogeros, Athanasios Ntalakas, M. Gergatsoulis et al.· 0 citations
The WMT26 General MT task evaluates systems on 10 language pairs that have no human references (neither translated from scratch nor post-edited from MT output by humans) and how the pseudo-references were built is described.
Diptesh Kanojia, Chi-Kiu Lo, Archchana Sindhujan et al.· 0 citations
Translation quality estimation (QE) must support both sentence-level scoring and local error diagnosis under limited, heterogeneous supervision. This study proposes MCA-LLM-QEER, which combines leakage-controlled multi-source corpus augmentation, source-language/crosslingual/ target-language semantic evidence, explicit...
Wei Wang, Wei-Feng Liu· International journal of pat...· 0 citations
Related blog posts
MIT News · Artificial Intelligence· news.mit.eduSep 24, 2026
A new method, called CW-Net, translates the reasoning process of an autonomous vehicle’s AI system into understandable concepts that explain its behavior.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.