Q1 How strong is it across many tests combined? AA Intelligence Index · Bottom 12%#104 Q1 How strong is it across many tests combined? AA Intelligence Index · Bottom 12%#104 Q2 Which AI gives the answers people prefer most? Arena Text · Bottom 9%#90 Q3 Does it have broad knowledge across expert fields? MMLU-Pro · Bottom 36%#19
Q1 Bottom 12%#104 of 117 11.6 Reasoning effort: High Q2 1352 Q3 80.8% By effort Low 10.2 High 11.6
About AA Intelligence Index More about AA Intelligence Index
Q1 Can it handle customer service by the rules? τ³-Banking · Bottom 27%#59 Q1 Can it handle customer service by the rules? τ³-Banking · Bottom 27%#59 Q2 Can it run a small business alone for a year and make money? Vending-Bench 2 · Bottom 7%#45 Q3 Can it handle customer service by the rules? TAU2 · Bottom 25%#61 Q4 Can it finish long tasks that take people hours? METR Time Horizon · Bottom 10%#10
Q1 Bottom 27%#59 of 79 12.8% Reasoning effort: High Q2 −$22 Q3 65.8% Q4 42 min By effort Low 2.9% High 12.8%
About τ³-Banking More about τ³-Banking
Q1 Can it pull scattered facts from several long documents into one answer? Long Context Reasoning · Bottom 13%#102 Q1 Can it pull scattered facts from several long documents into one answer? Long Context Reasoning · Bottom 13%#102 Q2 Which AI gives the business answers people prefer most? Arena Business, Management & Finance · Bottom 10%#89 Q3 Can it do professional work in banking, consulting and law? APEX-Agents · Bottom 3%#46 Q1 Bottom 13%#102 of 116 52.0% Reasoning effort: High Q2 1352 Q3 4.4% Q4 58.2% Q5 44.4% Q6 21.5% Q7 28.3% By effort Low 46.0% High 52.0%
About Long Context Reasoning More about Long Context Reasoning
Q1 Which AI gives answers people prefer most, without false claims? Arena Text Factuality · Bottom 6%#86 Q1 Which AI gives answers people prefer most, without false claims? Arena Text Factuality · Bottom 6%#86 Q2 Can it use search and code to solve problems that stump experts? Humanity's Last Exam · Bottom 3%#44 Q3 Does it avoid making things up when summarizing? Vectara Hallucination Rate · Bottom 9%#34 Q4 Does it avoid making things up when summarizing? Vectara Factual Consistency · Bottom 9%#34
Q1 Bottom 6%#86 of 90 1385 Q2 19.0% Q3 14.2% Q4 85.8% About Arena Text Factuality More about Arena Text Factuality
Q1 Can it operate a computer itself to finish hours-long development jobs? Terminal-Bench 4.0 · Bottom 27%#59 Q1 Can it operate a computer itself to finish hours-long development jobs? Terminal-Bench 4.0 · Bottom 27%#59 Q2 Which AI gives the coding answers people prefer most? Arena Coding · Bottom 8%#91 Q3 Can it find better answers to optimization problems with no single right answer? ALE-Bench · Bottom 15%#64 Q1 Bottom 27%#59 of 79 0.0% Reasoning effort: High Q2 1390 Q3 576 Q4 26.2% Q5 23.5% Q6 30.4 Q7 87.8% Q8 34.0% Q9 48.2% Q10 41.8% Q11 1385 Q12 18.7% Q13 1.41× Q14 11.0% About Terminal-Bench 4.0 More about Terminal-Bench 4.0
Q1 Which AI gives the writing and literature answers people prefer most? Arena Writing, Literature & Language · Bottom 8%#91 Q1 Which AI gives the writing and literature answers people prefer most? Arena Writing, Literature & Language · Bottom 8%#91 Q2 Which AI do people prefer most in long conversations? Arena Multi-turn · Bottom 9%#90 Q3 Does it follow tricky rules like format and length without missing any? IFBench · Bottom 50%#41 Q1 Bottom 8%#91 of 97 1309 Q2 1327 Q3 69.0% Q4 1278 Q5 1323 Q6 7.71 / 10 Q7 1286 About Arena Writing, Literature & Language More about Arena Writing, Literature & Language
Q1 Can it solve problems that stump experts across fields? Humanity's Last Exam · Bottom 28%#86 Q1 Can it solve problems that stump experts across fields? Humanity's Last Exam · Bottom 28%#86 Q2 Can it solve research-level physics problems? CritPt · Bottom 21%#63 Q3 Which AI gives the math answers people prefer most? Arena Math · Bottom 11%#85 Q1 Bottom 28%#86 of 117 19.6% Reasoning effort: High Q2 1.1% Q3 1381 Q4 78.2% Q5 88.9% Q6 93.4 Q7 93.4% Q8 25.0% Q9 1361 Q10 1384 Q11 22.1% Q12 20.0% Q13 2.0% Q14 76.3% Q15 22.1 By effort Low 5.9% High 19.6%
About Humanity's Last Exam More about Humanity's Last Exam
Q1 Which AI do people prefer most when asking in Korean? Arena Korean · Bottom 7%#75 Q1 Which AI do people prefer most when asking in Korean? Arena Korean · Bottom 7%#75 Q2 Which AI do people prefer most when asking in Japanese? Arena Japanese · Bottom 14%#66 Q3 Can it use expert knowledge in many languages? MMMLU · Bottom 6%#18 Q1 Bottom 7%#75 of 79 1266 Q2 1327 Q3 81.3% Q4 1362 Q5 1357 Q6 1367 Q7 1339 Q8 92.5% Q9 85.6% Q10 86.0% Q11 90.0% Q12 90.0% Q13 64.4% Q14 73.4% Q15 55.5% Q16 92.0% Q17 85.5% Q18 95.6% Q19 51.8% Q20 79.1% Q21 73.0% Q22 64.0% Q23 80.0% Q24 30.0% Q25 11.0% Q26 74.0% Q27 8.9% Q28 74.0% Q29 59.0% Q30 47.0% Q31 94.0% Q32 55.0% Q33 73.5% Q34 1.0% Q35 45.2% Q36 4.7% Q37 74.0% Q38 84.2% Q39 69.8% Q40 46.4% Q41 13.1% Q42 87.7% Q43 76.8% Q44 56.6% Q45 75.7% About Arena Korean More about Arena Korean