Skip to content
Preprint

OmnilingualGAIA2: Evaluating the Multilingual Gap in Frontier AI Agents

Aug 2026 · 2 citations · 57 references
Computer Science

TL;DR

OmnilingualGAIA2 is introduced, a machine-translated expansion of the GAIA2 agentic benchmark, covering ten target languages spanning five writing systems, paired with a localised and human-calibrated multilingual verifier, and it is argued that multilingual agentic evaluation must become a standard part of the reporting protocol for globally deployed agents.

Abstract

Agentic benchmarks aim to measure how well AI agents plan, search, execute, and recover within realistic multi-tool environments, but they are almost exclusively in English. As AI agents are globally deployed to a linguistically diverse user base, whether agentic competence measured in English transfers to other languages remains an open question. We introduce OmnilingualGAIA2, a machine-translated expansion (with partial human- expert validation) of the GAIA2 agentic benchmark, covering ten target languages spanning five writing systems, paired with a localised and human-calibrated multilingual verifier. Evaluating seven frontier and open-weight agents, we find a universal cross-lingual gap of 8.8-18.4 pass@3 points that is agent-asymmetric in magnitude, concentrates on tool-orchestration rather than quantitative reasoning, and does not close with model scale. A stratified error attribution decomposes the gap as predominantly model-driven (55%), with a bounded translation-contamination floor of only 6.4% of scenario-language pairs. Human-expert linguistic analysis further identifies morphological cue loss and amplified ambiguity as the primary failure mechanisms in non-Latin-script languages. Our results argue that multilingual agentic evaluation must become a standard part of the reporting protocol for globally deployed agents.

View source

Similar papers

BabelArena: A Large-Scale Multilingual Benchmark for LLM Agents

BabelFlow is introduced, a benchmark-general agentic workflow that adapts existing agent benchmarks to new languages by analyzing runtime dependencies, coordinating structure-preserving translation, and combining multi-layer verification with human review to preserve task and evaluation semantics.

Peng Kuang, Yu-Chun Fan, Jiang-Nan Li et al. · 0 citations
#artificial intelligence Preprint Sep 2026

WorldBench: Culturally Grounded Benchmark for Multilingual Agents

WorldBench is presented: a comprehensive, multilingual benchmark of genuine, persona-grounded everyday workflows, where agents can act in a sandbox via structured actions, and Constrained Task Success (CTS), which combines natural language instructions and testbeds to score task completion, minimal modification, and ot...

Leonardo Ranaldi, Sherrie Shen, Jushi Kai et al. · 0 citations
Preprint Aug 2026

From Atomic to Agentic: Towards Interpretable Evaluation of LLMs'Agentic Mathematical Capabilities

Experiments reveal that models with similar end-to-end accuracy can exhibit markedly different agentic capability profiles, demonstrating that process-level evaluation is crucial for interpreting the true potential of LLMs and guiding the development of next-generation mathematical agents.

Jiayi Kuang, Ying-Hui Li, Yun-Ze Song et al. · 0 citations
Preprint Aug 2026

An Actionable Diagnosis of Multilingual, Multi-Agent Planning Failures

To test whether the taxonomy supports mitigation, TART, Taxonomy-Guided Actionable Representation, is introduced that makes the taxonomy's key aspects explicit to the planner and downstream sub-agents and consistently improves performance.

Vikas Pahuja, J. Brokman, O. Hofman et al. · 0 citations
Preprint Aug 2026

AgentJudgeBench: A Multi-Difficulty Benchmark for Evaluating LLM Judges on Agentic Tool-Calling

The first benchmark to systematically study LLM-as-a-judge reliability for agentic tool-calling over workflow DAGs, as distinct from the broader LLM-as-a-judge task of open-ended text or preference evaluation, exposes fundamental limitations of current LLM judges and yields practical guidelines for reliable evaluation...

Abhigya Verma, Amit Kumar Saha, Seganrasan Subramanian et al. · 1 citation · ⚡1
Preprint Aug 2026

HopRefusalBench: Diagnosing Refusal Failures in Search-Augmented Agents for Multi-Hop Reasoning

This work introduces HopRefusalBench, the first controlled benchmark of refusal within multi-hop search, and proposes a final-outcome taxonomy spanning target-aware refusal, pseudo-refusal, hallucinated completion, and search-budget exhaustion, together with source-aware trajectory metrics for post-trigger continuation...

Jianan Xie, Xin Sun, Zhongqi Chen et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.