Applied AI
October 02, 2026
|
4 min read
Introducing AURORA, the Multilingual AI Leaderboard that measures Frontier AI on non-English agentic tasks
AURORA is LILT's multilingual AI leaderboard. It measures frontier models on non-English agentic tasks using localized benchmarks designed and verified by native-language experts, rather than translated tests. Results show model rankings flip by language and task: Opus 5.5 leads coding and customer support, Gemini 3.8 Flash leads multi-turn conversation.
LILT Team

LLM leaderboards show frontier models nearly neck-and-neck, making it look like general reasoning has been solved. But for AI teams shipping agents to users in Tokyo, Berlin, or Riyadh, there's a catch: those leaderboards almost entirely stop at English.
Nearly every published measure of frontier progress relies on English prompts or machine-translated test sets. That was fine when LLMs mostly summarized or translated text. Now agents write code and handle multi-turn customer support, and a failure outside English means a broken build, an abandoned cart, or a wrong operational decision.
The problem: why translated evals fail
Teams assume strong English performance carries over to other languages. Our data shows it doesn't. Machine translation introduces artifacts and strips out cultural context. Aggregate global scores mask a key fact: model rankings flip dramatically depending on the language and task.
Introducing AURORA: the multilingual AI leaderboard
To fix this industry gap, we are launching AURORA: LILT's Multilingual AI Leaderboard. AURORA stands for Agentic Understanding and Reasoning in Original-language, Real-world Assessments.
Rather than relying on translated tests, AURORA evaluates models using LILT's suite of localized, multilingual benchmarks, adapted directly from industry standards, to measure non-English performance on agentic, multimodal, and socio-cultural tasks. What sets these benchmarks apart is their creation: it features tasks that are designed and verified by native-language domain experts. These evaluations are specifically designed to reflect real-world enterprise use cases rooted in authentic language, regional context, and culture.
These LILT benchmarks include:
- Multilingual Terminal-Bench: Evaluates software development and coding capabilities across 10 languages, representing real-world challenges in software written for native-language users. Read the Terminal-Bench-LILT benchmark.
- Multilingual τ³-bench: Tests multi-turn agent interaction across various sectors, including banking, retail, telecom, and airline customer support.
- Multilingual MultiChallenge: Assesses contextual memory, self-coherence, and instruction-following within extended contexts.
- Multilingual GAIA-v2-LILT: Measures complex tool utilization and agentic reasoning capabilities. Read the GAIA-v2-LILT benchmark.
We recently completed evaluations against these benchmarks across both proprietary and open-weights state-of-the-art models. The sections below provide a comprehensive breakdown of our benchmark results and key findings.
Multilingual Terminal-Bench
Multilingual Terminal-Bench is our multilingual version of Terminal-Bench [1] and is intended to be a predictor of model performance on terminal-based tasks in 10 languages. The benchmark consists of 324 tasks, each designed in collaboration with a native speaker and contains a minimum of 30 tasks per language.
Key findings:
- Claude Opus 5.5 pulls out a big lead (~3 pp) over its already capable predecessor, Opus 5, at less than 60% of the price
- GPT-6 Sol edges ahead of GPT-5.6 Sol at less than half the cost
- Meta's latest model has comparatively low per-token price, but uses far more tokens and thus costs more per-task than more "expensive" models
- GLM 5.3 Flash (fp8) significantly outperforms GPT-5.4 at 8% of the price of the OpenAI model and is catching up to Opus 4.7
Opus 5.5 is the new leaderboard winner, showing consistent improvements over Opus 5 in task resolution for almost every language, with the biggest gains coming in Chinese and Korean, with a slight regression in German. While this model release brings a 20% reduction in token costs, usage analysis shows that this model also uses significantly fewer tokens (about 30% on average) to complete tasks resulting in an overall reduction of 45% from Opus 5.
GPT-6 Sol is the highest performing model from OpenAI, but only just. The last three models from OpenAI all score within 2 pp of each other, indicating a plateauing of multilingual coding capabilities. However, the pricing of this model (-50%) means it costs significantly less on coding tasks than before.
Google's Gemini 3.8 Flash ranks above GPT-6 Sol, by 2.5 pp. However, it requires more turns to get to the solution and ends up costing 7x that of GPT-6 Sol.
Meta's Muse Spark 1.3 trails slightly behind its predecessor on pass rate as it uses significantly more tokens and turns per-task and runs into time limits. Open-weights models are also catching up to the best proprietary models from ~6 months ago, with GLM 5.3 Flash significantly outperforming GPT-5.4 (+19.3 pp) at 8% of the cost of GPT-5.4, on average. Cohere's Command A+ scores 4.6%, with trajectory analysis showing that the model struggles with tool calls, input file resolution and script recognition.
| Model | Median turns per-task |
|---|---|
Opus 5.5 | 6 |
Opus 5 | 6 |
Gemini 3.8 Flash | 41 |
GPT-6 Sol | 5 |
GPT-5.6 Sol | 5 |
Muse Spark 1.3 | 11 |
GLM 5.3 Flash | 7 |
Muse Spark 1.1 | 8 |
Command A+ | 5 |
Multilingual τ³-bench
Multilingual τ³-bench is our adaptation of τ²-Bench [2] into German, Korean, Turkish and Farsi.
Key findings:
- Opus 5.5 and GPT-6 Sol both bring performance improvements and a reduction in token price over their predecessors, posting strong results on this benchmark while reducing per-task cost.
- Opus 5.5 takes top spot, with a ~9 pp gap to GPT-6
- Gemini 3.8 Flash brings Opus 5 performance at 1/3rd the cost per-task
Opus 5.5 wins comprehensively across all languages, showing gains in every non-English language in almost every domain over its predecessor. It also costs less than half of Opus 5 to solve tasks on average; $0.28 for Opus 5.5 vs $0.59 for Opus 5. Both Opus models show a tendency to switch languages during the conversation, but this is reduced by 80% for Opus 5.5.
It's a similar story with GPT-6 Sol, which improves over its predecessor across the board with very small regressions. The improvements are enough to get it to second place for the leaderboard, but it still trails Opus 5.5 by a significant amount. GPT-6 Sol also reduces price per-task compared to its predecessor; $0.16 vs $0.26 for GPT-5.6 Sol.
| Airline (Δ pp) | Retail (Δ pp) | Telecom (Δ pp) | Banking⁺ (Δ pp) | |
|---|---|---|---|---|
Opus 5.5 vs Opus 5 | ||||
Biggest gain (pass^4) | de (+6.0) | ko (+51.8) | ko (+51.7) | tr (+5.0) |
Biggest loss (pass^4) | en (-2.0) | n/a | n/a | en, ko (-5.0) |
GPT-6 Sol vs GPT-5.6 Sol | ||||
Biggest gain (pass^4) | en (+10.0) | fa (+28.0) | fa (+24.5) | ko (+5.0) |
Biggest loss (pass^4) | de (-2.0) | n/a | n/a | fa (-5.0) |
⁺ Banking domain is evaluated on a 20-task subset.
Google's Gemini 3.8 Flash performs comparably to Opus 5, but costs less than a third of Anthropic's model.
Cohere has a much stronger showing for this benchmark with Command A+, though it scores below other models tested. The largest drops in performance, compared to English, are in Korean and Persian, at 18.2 and 19.2 pp respectively.
These results show that for long conversations in customer support scenarios, models still have some way to go to be equally helpful in all languages.
Multilingual MultiChallenge
Multilingual MultiChallenge is our translated and human-verified version of the MultiChallenge [3] benchmark, which measures the performance of LLMs on multi-turn natural conversations.
Key findings:
- Gemini takes top spot on every non-English language
- GPT-6 Sol regresses on nearly every language compared to GPT-5.6 Sol
Gemini 3.8 Flash leads this benchmark for all non-English languages, followed by Opus 5.5. This is especially impressive considering that Gemini costs less than 1/6th per-task. OpenAI scores comparatively lower in this benchmark, showing a regression on nearly every language, except Arabic (+0.1 pp). The 4.4 pp regression in English stands out, as frontier models continue to ace English benchmarks.
Kimi-K3 and Cohere's Command A+ represent the open-weights models. K3 is a stronger performer in all languages, beating Opus 4.6 in all non-English languages. Command A+ trails other models tested in all languages except Arabic, where it performed better than GPT-5.2.
Explore the full results
No single model leads across every language, task, and cost profile. The right choice depends on where your users are and what your agents need to do. The full AURORA results, including methodology, per-language scores, per-task costs, and domain-level breakdowns for each benchmark, are available at aurora.lilt.com.
We will keep updating the leaderboard as new models are released and new languages are added. Follow us on LinkedIn or X or contact us.
References
[1] Merrill, Mike A. et al. "Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces." ArXiv abs/2601.11868 (2026): n. pag.
[2] Barres, Victor, et al. "τ²-Bench: Evaluating Conversational Agents in a Dual-Control Environment." arXiv preprint arXiv:2506.07982 (2025).
[3] Deshpande, Kaustubh, et al. "MultiChallenge: A realistic multi-turn conversation evaluation benchmark challenging to frontier LLMs." Findings of the Association for Computational Linguistics: ACL 2025. 2025.
Frequently Asked Questions
What is AURORA?
AURORA is LILT's Multilingual AI Leaderboard. The name stands for Agentic Understanding and Reasoning in Original-language, Real-world Assessments. It measures how frontier AI models perform on non-English agentic, multimodal, and socio-cultural tasks.
Why do English-only or translated AI benchmarks fall short?
Most published benchmarks rely on English prompts or machine-translated test sets. Machine translation introduces artifacts and strips out cultural context. Model rankings also flip depending on the language and task, so aggregate global scores hide real differences.
How is AURORA different from other LLM leaderboards?
AURORA doesn't rely on translated tests. It uses LILT's localized, multilingual benchmarks, adapted from industry standards. Native-language domain experts design and verify the tasks, which reflect real enterprise use cases.
Which benchmarks does AURORA include?
AURORA includes four benchmarks: Multilingual Terminal-Bench (software development and coding across 10 languages); Multilingual τ³-bench (multi-turn agent interaction in banking, retail, telecom, and airline customer support); Multilingual MultiChallenge (contextual memory, self-coherence, and instruction-following in extended contexts); and Multilingual GAIA-v2-LILT (complex tool use and agentic reasoning). The leaderboard is updated as new models are released and new languages are added.
Which AI model performs best on multilingual coding tasks?
Opus 5.5 is the leaderboard winner on Multilingual Terminal-Bench. It leads its predecessor, Opus 5, by about 3 percentage points. Its biggest gains are in Chinese and Korean, with a slight regression in German. The last three OpenAI models score within 2 points of each other, which suggests multilingual coding performance is plateauing.
Which model is best for multilingual customer support agents?
Opus 5.5 takes the top spot on Multilingual τ³-bench, about 9 points ahead of GPT-6 Sol. It costs $0.28 per task versus $0.59 for Opus 5. Its tendency to switch languages mid-conversation dropped by 80% compared with Opus 5.
Which model leads on multi-turn conversation in non-English languages?
Gemini 3.8 Flash leads Multilingual MultiChallenge in every non-English language, followed by Opus 5.5. It does so at less than one-fifth of the per-task cost of the other leading models.
How many tasks and languages does Multilingual Terminal-Bench cover?
It has 324 tasks across 10 languages, with at least 30 tasks per language. Each task was designed in collaboration with a native speaker.
Which languages does Multilingual τ³-bench cover?
German, Korean, Turkish, and Farsi, adapted from τ²-Bench.
Do open-weights models compete with proprietary models on multilingual tasks?
They are catching up. GLM 5.3 Flash (fp8) beats GPT-5.4 by 19.3 points on Terminal-Bench at about 8% of the cost. Open-weights models now match proprietary models from about six months ago. Kimi-K3 beats Opus 4.6 in all non-English languages on MultiChallenge. Cohere's Command A+ trails the other models tested, though it does better on τ³-bench.
Does the cheapest model per token cost the least per task?
No. Meta's latest model has a low per-token price but uses far more tokens and turns, so it costs more per task than some pricier models. Gemini 3.8 Flash ranks above GPT-6 Sol on Terminal-Bench by 2.5 points but takes more turns and costs about 7x as much per task.
Do AI models perform equally well in every language?
No. Compared with English, Command A+ drops 18.2 points in Korean and 19.2 points in Persian on τ³-bench. For long customer support conversations, models still have some way to go before they are equally helpful in all languages.
Where can I see the full AURORA results?
Methodology, per-language scores, per-task costs, and domain-level breakdowns are at aurora.lilt.com. LILT will keep updating the leaderboard as new models and languages are added. You can reach the team at lilt.com/contact.
Request Dataset Samples
English scores don't tell the whole story. Let's measure yours. Evaluate models and agents on tasks grounded in language, culture, and region.
Book a MeetingShare this post
Request Dataset Samples
English scores don't tell the whole story. Let's measure yours. Evaluate models and agents on tasks grounded in language, culture, and region.
Book a MeetingShare this post