Skip to content

A Systematic Comparison of Multilingual Interpretability Methods Reveals Anisotropy-Driven Failures

Sep 2026 · 0 citations · 46 references
Computer Science

TL;DR

Only ILO's correlation with cross-lingual transfer (Spearman's $\rho = 0.90$) survives controls for model size, family, and per-task variation and is recommended as the primary sharing metric to be reported alongside anisotropy diagnostics.

Abstract

Multilingual language models develop shared cross-lingual representations, and various interpretability methods claim to quantify this sharing. These methods have been developed largely in isolation, and when they disagree, it is unclear whether the disagreement reflects a property of the model or an artifact of the measurement. We compare four sharing metrics (CKA, ANC, GMM dominance per token, and ILO) across 21 base models from five families (125M-14B parameters) and correlate each with cross-lingual transfer on five downstream tasks. We find that the metrics differ in their quantification of cross-lingual sharing in these models and suggest that the disagreement traces to anisotropy, the tendency of representations to cluster in a narrow cone of the embedding space. Only ILO's correlation with cross-lingual transfer (Spearman's $\rho = 0.90$) survives controls for model size, family, and per-task variation. We therefore recommend ILO as the primary sharing metric, to be reported alongside anisotropy diagnostics.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

A Systematic Evaluation of Cross-Lingual Consistency Enhancement Methods in Multilingual Language Models

Multilingual language models often produce inconsistent answers to semantically equivalent questions across languages, motivating methods to improve cross-lingual consistency (CLC). However, existing methods are typically evaluated using different models, tasks, and protocols, leaving their relative strengths unclear....

Jirui Qi, Ming-Yang Wang, Hinrich Schütze et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Lingua Franca or Probing Artifact? Rethinking Latent Language in Multilingual LLMs

This work finds that identification probes systematically disagree: the GMM-based representation probe, which draws evidence from hidden state geometry, shows earlier cross-lingual mixing, whereas decoding-based probes, which rely on output-space decodability, retain sharper language-specific and more English-biased si...

Deniz Bayazit, Badr AlKhamissi, Antoine Bosselut · 0 citations
#artificial intelligence Preprint Sep 2026

Why Pretraining Fails to Share Cross-Lingual Knowledge

Large Language Models (LLMs) have made remarkable progress in the processing and modeling of many languages. Yet, unlike human multilinguals, they exhibit surprisingly limited cross-lingual knowledge transfer. While this limitation is well documented, its origins during multilingual training remain unclear. We pretrain...

Adam Gaber, Uriel Dolev, Elisabeth Fittschen et al. · 0 citations
Preprint Aug 2026

Cross-Lingual Alignment Without Joint Training: Do Monolingual Language Models Converge on Universal Representations?

It is confirmed that cross-lingual alignment can emerge from the structure of language and the information it carries rather than from joint training, and this points to practical future directions including model stitching, merging, and modular multilingual systems built from monolingual components.

Ej Zhou, Suchir Salhan, Catherine Arnett et al. · 1 citation
Preprint Aug 2026

Divergent large language model predictions from convergent representations in ambiguous word pairs

This work investigates how decoder-only transformers resolve lexical ambiguity through layer-by-layer analysis of three models spanning three parameter sizes, finding that representations become maximally distinct in middle layers, then partially reconverge in late layers, while the KL divergence between their next-tok...

K. Scott, Narun Pat, Veronica Liesaputra · 1 citation

Related blog posts

MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.