Skip to content

BabelArena: A Large-Scale Multilingual Benchmark for LLM Agents

Sep 2026 · 0 citations · 22 references
Computer Science

TL;DR

BabelFlow is introduced, a benchmark-general agentic workflow that adapts existing agent benchmarks to new languages by analyzing runtime dependencies, coordinating structure-preserving translation, and combining multi-layer verification with human review to preserve task and evaluation semantics.

Abstract

Large language model (LLM) agents increasingly execute multi-step workflows through tool use and interaction with users and environments. However, current agent evaluations are largely English-centric, limiting our understanding of agent capabilities in multilingual settings. We introduce BabelFlow, a benchmark-general agentic workflow that adapts existing agent benchmarks to new languages by analyzing runtime dependencies, coordinating structure-preserving translation, and combining multi-layer verification with human review to preserve task and evaluation semantics. Using BabelFlow, we construct BabelArena, a task-aligned benchmark comprising 16,146 instances derived from 702 canonical tasks across four benchmark families, 13 domains, and 23 languages. Experiments with five frontier models show that no single model dominates across benchmark families and that cross-language disparities extend well beyond task success. Lower-resource languages exhibit distinct failure patterns, with larger shares of tool-use and control-flow errors rather than answer-quality errors alone, pointing to gaps in reliable task execution across the resource levels of these languages. On the same tasks, agents in low-resource languages also consume substantially more tokens than in English (up to roughly twice the input) without proportional increases in interaction length, and language consistency degrades further on tasks requiring structured output, where switches are directed overwhelmingly toward English. We believe BabelArena provides a foundation for advancing research on reliable and efficient multilingual agents.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

WorldBench: Culturally Grounded Benchmark for Multilingual Agents

WorldBench is presented: a comprehensive, multilingual benchmark of genuine, persona-grounded everyday workflows, where agents can act in a sandbox via structured actions, and Constrained Task Success (CTS), which combines natural language instructions and testbeds to score task completion, minimal modification, and ot...

Leonardo Ranaldi, Sherrie Shen, Jushi Kai et al. · 0 citations
Preprint Aug 2026

OmnilingualGAIA2: Evaluating the Multilingual Gap in Frontier AI Agents

OmnilingualGAIA2 is introduced, a machine-translated expansion of the GAIA2 agentic benchmark, covering ten target languages spanning five writing systems, paired with a localised and human-calibrated multilingual verifier, and it is argued that multilingual agentic evaluation must become a standard part of the reporti...

Andrea Caciolai, P. L. Cabot, Chierh Cheng et al. · 2 citations
Conference Open access Sep 2026

A Review on Test-Time Scaling for Agentic Large Language Models

A novel RAIE taxonomy along four scaling dimensions is proposed, which optimizes the entire thought process through search algorithms and self-verification, and introduces a task-oriented guideline for choosing the best TTS strategy.

Jia-Yu An, Zheng Chen, Yong-Cheng Jing et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Exploring Collaboration between a language and a non-language agent

To solve LLM collaboration with non-language agents, latent state internalization is introduced, which projects the subagent's continuous representations directly into the LLM's token stream as learned state tokens, with dynamic re-encoding as actions advance the environment state.

Harini S.I., Somesh Singh, Yaman Kumar Singla et al. · 1 citation
Preprint Aug 2026

An Actionable Diagnosis of Multilingual, Multi-Agent Planning Failures

To test whether the taxonomy supports mitigation, TART, Taxonomy-Guided Actionable Representation, is introduced that makes the taxonomy's key aspects explicit to the planner and downstream sub-agents and consistently improves performance.

Vikas Pahuja, J. Brokman, O. Hofman et al. · 0 citations
#artificial intelligence Preprint Sep 2026

DAREBench: Deployment-Aware and Reliable Evaluation of Models as Agents

DAREBench (Deployment-Aware and Reliable Evaluation of Models as Agents), a benchmark designed to capture workload variation and support reliable agent evaluation, is introduced, suggesting that agent deployment and model selection should consider workload profiles, deployment mode, and accuracy--cost trade-offs rather...

Yu Liu, Zhi-Lin Liu, Zhi-Wei Yang et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.