Q1 How strong is it across many tests combined? AA Intelligence Index · Top 21%#24 Q1 How strong is it across many tests combined? AA Intelligence Index · Top 21%#24 Q2 Can it solve brand-new problems, not just memorized ones? LiveBench Overall · Top 46%#23
Q1 Top 21%#24 of 117 39.8 Q2 76.2 About AA Intelligence Index More about AA Intelligence Index
Q Can it handle customer service by the rules? τ³-Banking · Top 11%#8
Q Can it handle customer service by the rules? τ³-Banking · Top 11%#8
Top 11%#8 of 79 45.4% About τ³-Banking More about τ³-Banking
Q1 Can it pull scattered facts from several long documents into one answer? Long Context Reasoning · Top 32%#36 Q1 Can it pull scattered facts from several long documents into one answer? Long Context Reasoning · Top 32%#36 Q2 Can it read tables and reshape them as needed? LiveBench Data Analysis · Bottom 50%#26 Q3 Can it do real work from many different jobs? JobBench · Top 39%#10 Q4 Can it read business PDFs and answer accurately? GDP.pdf · Bottom 34%#23
Q1 Top 32%#36 of 116 79.7% Q2 74.2 Q3 55.7% Q4 16.6% About Long Context Reasoning More about Long Context Reasoning
Q1 Can it operate a computer itself to finish hours-long development jobs? Terminal-Bench 4.0 · Top 27%#21 Q1 Can it operate a computer itself to finish hours-long development jobs? Terminal-Bench 4.0 · Top 27%#21 Q2 Can it run and fix code until the task is done? LiveBench Agentic Coding · Top 16%#8 Q3 Can it fix real issues in real software projects? SWE-Bench Pro · Top 25%#13 Q1 Top 27%#21 of 79 25.3% Q2 61.6 Q3 62.5% Q4 86.1% Q5 73.1 Q6 50.6% Q7 72.5 About Terminal-Bench 4.0 More about Terminal-Bench 4.0
Q1 Does it have a good feel for language, like word puzzles and spelling? LiveBench Language · Bottom 22%#40 Q1 Does it have a good feel for language, like word puzzles and spelling? LiveBench Language · Bottom 22%#40 Q2 Does it follow tricky rules like format and length without missing any? LiveBench Instruction Following · Top 10%#5
Q1 Bottom 22%#40 of 50 74.6 Q2 77.1 About LiveBench Language More about LiveBench Language
Q1 Can it solve problems that stump experts across fields? Humanity's Last Exam · Top 38%#44 Q1 Can it solve problems that stump experts across fields? Humanity's Last Exam · Top 38%#44 Q2 Can it solve graduate-level science problems? GPQA Diamond · Top 18%#18 Q3 Can it solve competition-level math problems? LiveBench Math · Bottom 26%#38 Q4 Can it solve logic puzzles? LiveBench Reasoning · Top 44%#22
Q1 Top 38%#44 of 117 38.0% Q2 92.3% Q3 85.8 Q4 87.4 About Humanity's Last Exam More about Humanity's Last Exam
Q Can it read charts and answer accurately? CharXiv · Top 10%#3
Q Can it read charts and answer accurately? CharXiv · Top 10%#3
Top 10%#3 of 31 90.6% Self-reported by developer About CharXiv More about CharXiv
Q1 Can it sort customer questions into the right categories? Jevals Banking77 · Bottom 34%#5 Q1 Can it sort customer questions into the right categories? Jevals Banking77 · Bottom 34%#5 Q2 Can it read a medical paper and answer yes or no correctly? Jevals PubMedQA · Top 50%#3 Q3 Can it correctly score how helpful an AI answer is? Jevals HelpSteer2 · Bottom 50%#4
Q1 Bottom 34%#5 of 6 62.4 Q2 62.4 Q3 -1.37 About Jevals Banking77 More about Jevals Banking77
Q1 Japanese MT-Bench Japanese MT-Bench · Top 26%#16 Q1 Japanese MT-Bench Japanese MT-Bench · Top 26%#16 Q2 jaster (0-shot) jaster · Top 39%#24 Q3 Nejumi LLM Leaderboard 4 Nejumi LLM Leaderboard 4 · Top 25%#15 Q1 Top 26%#16 of 62 97.2% Reasoning effort: Extra High Q2 84.7% Q3 82.9% Q4 87.5% Q5 81.0% About Japanese MT-Bench More about Japanese MT-Bench