Skip to content

EviSI: An Evidence-Based Evaluation Agent for Simultaneous Interpreting

Sep 2026 · 0 citations · 30 references
Computer Science

TL;DR

EviSI, a large language model evaluation agent combining Multidimensional Quality Metrics (MQM) with criteria developed with professional interpreters, is proposed, showing positive system ranking correlations with COMET throughout.

Abstract

Low-latency simultaneous speech-to-speech translation must keep pace with ongoing speech while preserving key information. To meet these demands, systems use segmentation, reformulation and condensation to reorganize and rephrase information. However, metrics developed for text translation, including BLEU and COMET, may not consistently distinguish faithful adaptations from semantic errors. We propose EviSI, a large language model evaluation agent combining Multidimensional Quality Metrics (MQM) with criteria developed with professional interpreters. Shared source evidence guides assessment across four dimensions: Anchor, Event, Logic and Fluency. Verified errors are deduplicated before deterministic scoring. On human-rated English to Chinese and Chinese to English data, EviSI recovers the aggregate English to Chinese human system ranking. Mean within-dataset Kendall correlations for system rankings reach 0.707 and 0.467, respectively, exceeding evaluated BLEU and COMET baselines. A multilingual extension to five directions without human ratings retains the dimensions and scoring rule, showing positive system ranking correlations with COMET throughout.

View source

Similar papers

#natural language process... Preprint Sep 2026

SEAR: Segment-Evidence-Aware Routing for Weak-to-Strong Multilingual Speech MCQ

This paper describes the system for Task~2 of the second Multilingual Conversational Speech Language Model (MLC-SLM) Challenge, which adapt Qwen3-Omni-30B-A3B-Instruct with a segment-evidence-aware data and post-training pipeline and obtains 90.92% accuracy on the final official evaluation set.

H. Le, L. Nguyen, Minh Tri Dao · 1 citation
#natural language process... Preprint Sep 2026

Choosing the Right Language Mode at Inference Time for Multilingual Reliability

Multilingual large language models often struggle to reason in low- to mid-resource languages. Prior work has shown that translation can improve multilingual reasoning by helping models access stronger English-centric representations. This raises a central question: How much translation is needed for multilingual large...

Ekata Mitra, Ameeta Agrawal · 0 citations
#artificial intelligence Preprint Sep 2026

Beyond BLEU: A Case for Redefining Sign Language Translation Benchmarks

This work introduces an alternative inspired by language-learning assessment, using an open-weight-LLM QA protocol that measures salient content preservation that aligns more closely with human rankings and is six to seven times more paraphrase-invariant than BLEU-4.

Oline Ranum, Edward Fish, Simon Hadfield et al. · 0 citations
#natural language process... Preprint Sep 2026

mu-bench: A Multilingual Utterance Transcription Benchmark

Voice agents depend on accurate automatic speech recognition (ASR) to act on what callers say, yet ASR is evaluated on read, English-centric speech with word error rate (WER), which penalizes surface rather than semantic differences. We introduce mu-bench, a dataset of 4,270 caller utterances from 250 phone calls to an...

Andrea Li, Soham Ray · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.