Skip to content

Apples to Apples? Towards Comparable Crosslingual Language Model Evaluation

Aug 2026 · 0 citations · 85 references
Computer Science

TL;DR

The results show that several widely used normalized metrics introduce crosslinguistic biases rooted in tokenization, encoding, and orthographic differences, and sentence-level negative log-likelihood computed over semantically equivalent sequences provides more meaningful and consistent crosslingual comparisons.

Abstract

Crosslingual evaluation of language models that enables fair comparisons remains a fundamental challenge in multilingual NLP. Existing studies adopt a variety of downstream tasks and intrinsic metrics with different theoretical justifications, yet there has been little empirical investigation into whether these approaches yield meaningful crosslingual conclusions. We systematically examine crosslingual evaluation approaches using controlled monolingual language models trained on parallel data with varying tokenizer vocabulary sizes and model sizes, and further validate our findings on multilingual LLMs. We further discuss challenges in achieving comparable downstream evaluation across languages. Our results show that several widely used normalized metrics introduce crosslinguistic biases rooted in tokenization, encoding, and orthographic differences. In contrast, sentence-level negative log-likelihood computed over semantically equivalent sequences provides more meaningful and consistent crosslingual comparisons.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

A Systematic Evaluation of Cross-Lingual Consistency Enhancement Methods in Multilingual Language Models

Multilingual language models often produce inconsistent answers to semantically equivalent questions across languages, motivating methods to improve cross-lingual consistency (CLC). However, existing methods are typically evaluated using different models, tasks, and protocols, leaving their relative strengths unclear....

Jirui Qi, Ming-Yang Wang, Hinrich Schütze et al. · 0 citations
#natural language process... Preprint Sep 2026

A Systematic Comparison of Multilingual Interpretability Methods Reveals Anisotropy-Driven Failures

Only ILO's correlation with cross-lingual transfer (Spearman's $\rho = 0.90$) survives controls for model size, family, and per-task variation and is recommended as the primary sharing metric to be reported alongside anisotropy diagnostics.

Oskar Holmström, Marcel Bollmann, Marco Kuhlmann · 0 citations
Preprint Aug 2026

Predicting Multilingual Classification and Translation Performance of LLMs with Cross-Lingual Alignment -- Is English Enough?

A PMI-based translation metric is proposed, which is less dependent on the target language and correlates strongly with chrF, and finds that CLA with English predicts translation quality comparably to or better than source-target CLA.

Adnan Al Ali, Kathy Hämmerl, Jindrich Libovický et al. · 0 citations
2026

Beyond Literal Meaning: How LLMs Interpret Yemeni Proverbs

Results show that instruction-tuned models like GPT-4o and Gemini 1.5 Pro outperform smaller models in both automatic and human evaluations, and LLM-as-a-Judge evaluation correlates strongly with human assessment.

Nasser Thmer, Ali Allaith, Muhammad Shoaib · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.