Q1 How strong is it across many tests combined? AA Intelligence Index · Bottom 17%#99 Q1 How strong is it across many tests combined? AA Intelligence Index · Bottom 17%#99 Q2 Which AI gives the answers people prefer most? Arena Text · Bottom 7%#92 Q3 Does it have broad knowledge across expert fields? MMLU-Pro · Bottom 43%#17 Q4 Does it have broad knowledge across expert fields? MMLU ·
Q1 Bottom 17%#99 of 117 13.4 Reasoning effort: High Q2 1335 Q3 81.8% Q4 77.0% By effort None 8.70 Low 11.8 Medium 12.5 High 13.4
About AA Intelligence Index More about AA Intelligence Index
Q Can it handle customer service by the rules? TAU2 · Bottom 32%#55
Q Can it handle customer service by the rules? TAU2 · Bottom 32%#55
Bottom 32%#55 of 79 75.7% Reasoning effort: Medium By effort None 62.0% Low 71.9% Medium 75.7% High 72.8%
About TAU2 More about TAU2
Q1 Can it pull scattered facts from several long documents into one answer? Long Context Reasoning · Bottom 15%#100 Q1 Can it pull scattered facts from several long documents into one answer? Long Context Reasoning · Bottom 15%#100 Q2 Which AI gives the business answers people prefer most? Arena Business, Management & Finance · Bottom 9%#90
Q1 Bottom 15%#100 of 116 60.3% Reasoning effort: High Q2 1351 By effort None 18.7% Low 54.3% Medium 60.0% High 60.3%
About Long Context Reasoning More about Long Context Reasoning
Q Which AI gives answers people prefer most, without false claims? Arena Text Factuality · Bottom 5%#87
Q Which AI gives answers people prefer most, without false claims? Arena Text Factuality · Bottom 5%#87
Bottom 5%#87 of 90 1380 About Arena Text Factuality More about Arena Text Factuality
Q1 Which AI gives the coding answers people prefer most? Arena Coding · Bottom 9%#90 Q1 Which AI gives the coding answers people prefer most? Arena Coding · Bottom 9%#90 Q2 Can it find better answers to optimization problems with no single right answer? ALE-Bench · Bottom 5%#72 Q3 Can it operate a computer itself to finish development jobs on its own? Terminal-Bench 2.1 · Bottom 10%#74 Q1 Bottom 9%#90 of 97 1393 Q2 236 Q3 16.1% Q4 17.4% Q5 23.0 Q6 71.1% Q7 24.0% Q8 1386 About Arena Coding More about Arena Coding
Q1 Which AI gives the writing and literature answers people prefer most? Arena Writing, Literature & Language · Bottom 5%#94 Q1 Which AI gives the writing and literature answers people prefer most? Arena Writing, Literature & Language · Bottom 5%#94 Q2 Which AI do people prefer most in long conversations? Arena Multi-turn · Bottom 6%#93 Q3 Does it follow tricky rules like format and length without missing any? IFBench · Top 49%#39 Q1 Bottom 5%#94 of 97 1298 Q2 1323 Q3 70.7% Q4 1269 Q5 1327 Q6 1283 About Arena Writing, Literature & Language More about Arena Writing, Literature & Language
Q1 Can it solve problems that stump experts across fields? Humanity's Last Exam · Bottom 15%#101 Q1 Can it solve problems that stump experts across fields? Humanity's Last Exam · Bottom 15%#101 Q2 Which AI gives the math answers people prefer most? Arena Math · Bottom 6%#90 Q3 Can it solve graduate-level science problems? GPQA Diamond · Bottom 30%#74 Q1 Bottom 15%#101 of 117 11.6% Reasoning effort: High Q2 1332 Q3 81.1% Q4 94.3 Q5 94.3% Q6 1357 Q7 1347 By effort None 2.9% Low 4.0% Medium 9.0% High 11.6%
About Humanity's Last Exam More about Humanity's Last Exam
Q Can it solve college-level problems with images and diagrams? MMMU-Pro · Bottom 6%#36
Q Can it solve college-level problems with images and diagrams? MMMU-Pro · Bottom 6%#36
Bottom 6%#36 of 37 61.8% Self-reported by developer About MMMU-Pro More about MMMU-Pro
Q1 Which AI do people prefer most when asking in Korean? Arena Korean · Bottom 3%#78 Q1 Which AI do people prefer most when asking in Korean? Arena Korean · Bottom 3%#78 Q2 Which AI do people prefer most when asking in Japanese? Arena Japanese · Bottom 4%#73 Q3 Which AI do people prefer most on tough questions? Arena Hard Prompts · Bottom 8%#91 Q1 Bottom 3%#78 of 79 1246 Q2 1248 Q3 1357 Q4 1352 Q5 1360 Q6 1335 About Arena Korean More about Arena Korean