Q1 How strong is it across many tests combined? AA Intelligence Index · Bottom 11%#106 Q1 How strong is it across many tests combined? AA Intelligence Index · Bottom 11%#106 Q2 Does it have broad knowledge across expert fields? MMLU-Pro · Bottom 11%#26
Q1 Bottom 11%#106 of 117 11.3 Reasoning effort: None Q2 74.3% About AA Intelligence Index More about AA Intelligence Index
Q Can it handle customer service by the rules? TAU2 · Bottom 21%#64
Q Can it handle customer service by the rules? TAU2 · Bottom 21%#64
Bottom 21%#64 of 79 63.7% Reasoning effort: None About TAU2 More about TAU2
Q1 Can it pull scattered facts from several long documents into one answer? Long Context Reasoning · Bottom 6%#111 Q1 Can it pull scattered facts from several long documents into one answer? Long Context Reasoning · Bottom 6%#111 Q2 Can it read a contract hundreds of pages long and find exactly what's needed? CorpFin v2 · Bottom 7%#60 Q3 Can it read hospital records and assign the right diagnosis codes? MedCode · Bottom 6%#69 Q1 Bottom 6%#111 of 116 31.3% Reasoning effort: None Q2 52.5% Q3 28.3% Q4 44.4% Q5 37.9% About Long Context Reasoning More about Long Context Reasoning
Q1 Does it avoid making things up when summarizing? Vectara Hallucination Rate · Bottom 6%#35 Q1 Does it avoid making things up when summarizing? Vectara Hallucination Rate · Bottom 6%#35 Q2 Does it avoid making things up when summarizing? Vectara Factual Consistency · Bottom 6%#35 Q3 How well can it predict what will happen? ForecastBench Brier Index · Bottom 12%#25
Q1 Bottom 6%#35 of 36 17.8% Q2 82.2% Q3 56.7 About Vectara Hallucination Rate More about Vectara Hallucination Rate
Q1 Can it operate a computer itself to crack the toughest development jobs? TerminalBench Hard · Bottom 14%#69 Q1 Can it operate a computer itself to crack the toughest development jobs? TerminalBench Hard · Bottom 14%#69 Q2 How strong is its overall coding? AA Coding Index · Bottom 10%#91 Q3 Can it solve newly released programming problems? LiveCodeBench · Bottom 18%#25 Q4 Can it write code for scientific research? SciCode · Bottom 8%#109
Q1 Bottom 14%#69 of 79 14.4% Reasoning effort: None Q2 19.5 Q3 39.9% Q4 29.6% About TerminalBench Hard More about TerminalBench Hard
Q1 Does it follow tricky rules like format and length without missing any? IFBench · Bottom 2%#81 Q1 Does it follow tricky rules like format and length without missing any? IFBench · Bottom 2%#81 Q2 Can it turn a doctor–patient conversation into an accurate medical record? MedScribe · Bottom 32%#51
Q1 Bottom 2%#81 of 81 36.5% Reasoning effort: None Q2 77.5% About IFBench More about IFBench
Q1 Can it solve problems that stump experts across fields? Humanity's Last Exam · Bottom 6%#111 Q1 Can it solve problems that stump experts across fields? Humanity's Last Exam · Bottom 6%#111 Q2 Can it solve graduate-level science problems? GPQA Diamond · Bottom 6%#98 Q3 How strong is its overall math? AA Math Index · Bottom 13%#22 Q1 Bottom 6%#111 of 117 5.1% Reasoning effort: None Q2 63.7% Q3 34.3 Q4 34.3% Q5 56.0% About Humanity's Last Exam More about Humanity's Last Exam
Q1 HAE-RAE Bench v1 (Reading) HAE-RAE Bench v1 · Bottom 9%#45 Q1 HAE-RAE Bench v1 (Reading) HAE-RAE Bench v1 · Bottom 9%#45 Q2 Horangi 4 (Korean LLM Leaderboard) Horangi 4 · Bottom 5%#48 Q3 Horangi 4 — ALT (alignment) Horangi 4 — ALT · Bottom 25%#37 Q1 Bottom 9%#45 of 48 81.0% Q2 60.3% Q3 73.0% Q4 47.7% Q5 62.0% Q6 74.9% Q7 67.0% Q8 53.0% Q9 23.3% Q10 4.0% Q11 9.0% Q12 74.0% Q13 8.4% Q14 74.0% Q15 46.0% Q16 24.0% Q17 88.0% Q18 70.0% About HAE-RAE Bench v1 More about HAE-RAE Bench v1