Q How strong is it across many tests combined? AA Intelligence Index · Bottom 25%#89
Q How strong is it across many tests combined? AA Intelligence Index · Bottom 25%#89
Bottom 25%#89 of 117 18.2 Reasoning effort: Thinking By effort None 15.2 Thinking 18.2
About AA Intelligence Index More about AA Intelligence Index
Q1 Can it handle customer service by the rules? τ³-Banking · Bottom 14%#69 Q1 Can it handle customer service by the rules? τ³-Banking · Bottom 14%#69 Q2 Can it handle customer service by the rules? TAU2 · Top 22%#17 Q3 Can it pick the right tools and use them properly? MCP-Atlas · Bottom 30%#27
Q1 Bottom 14%#69 of 79 9.3% Reasoning effort: Thinking Q2 95.3% Q3 62.8% By effort None 5.4% Thinking 9.3%
About τ³-Banking More about τ³-Banking
Q Can it pull scattered facts from several long documents into one answer? Long Context Reasoning · Bottom 32%#81
Q Can it pull scattered facts from several long documents into one answer? Long Context Reasoning · Bottom 32%#81
Bottom 32%#81 of 116 71.7% Reasoning effort: Thinking By effort None 64.3% Thinking 71.7%
About Long Context Reasoning More about Long Context Reasoning
Q1 Can it operate a computer itself to finish hours-long development jobs? Terminal-Bench 4.0 · Bottom 27%#59 Q1 Can it operate a computer itself to finish hours-long development jobs? Terminal-Bench 4.0 · Bottom 27%#59 Q2 Can it fix real issues in real software projects? SWE-Bench Pro · Bottom 12%#47 Q3 Can it operate a computer itself to finish development jobs on its own? Terminal-Bench 2.1 · Bottom 33%#56 Q1 Bottom 27%#59 of 79 0.0% Reasoning effort: Thinking Q2 49.5% Q3 44.9% Q4 34.8% Q5 41.9 Q6 36.6% Q7 34.5% Q8 23.0% About Terminal-Bench 4.0 More about Terminal-Bench 4.0
Q Does it follow tricky rules like format and length without missing any? IFBench · Bottom 42%#48
Q Does it follow tricky rules like format and length without missing any? IFBench · Bottom 42%#48
Bottom 42%#48 of 81 64.4% Reasoning effort: Thinking By effort None 36.2% Thinking 64.4%
About IFBench More about IFBench
Q1 Can it solve problems that stump experts across fields? Humanity's Last Exam · Bottom 33%#80 Q1 Can it solve problems that stump experts across fields? Humanity's Last Exam · Bottom 33%#80 Q2 Can it solve research-level physics problems? CritPt · Bottom 12%#70 Q3 Can it solve graduate-level science problems? GPQA Diamond · Bottom 40%#63 Q1 Bottom 33%#80 of 117 22.2% Reasoning effort: Thinking Q2 0.3% Q3 84.1% Q4 86.7% Q5 20.4% Q6 78.9% Q7 44.4% Q8 26.0% Q9 22.0% Q10 73.9% Q11 29.7 By effort None 13.9% Thinking 22.2%
About Humanity's Last Exam More about Humanity's Last Exam
Q1 Can it read charts and answer accurately? CharXiv · Bottom 33%#22 Q1 Can it read charts and answer accurately? CharXiv · Bottom 33%#22 Q2 Can it solve college-level problems with images and diagrams? MMMU-Pro · Bottom 38%#24
Q1 Bottom 33%#22 of 31 78.0% Self-reported by developer Q2 75.3% About CharXiv More about CharXiv
Q1 Japanese MT-Bench Japanese MT-Bench · Bottom 39%#39 Q1 Japanese MT-Bench Japanese MT-Bench · Bottom 39%#39 Q2 jaster (0-shot) jaster · Bottom 31%#44 Q3 MRCR v2 (8-needle) MRCR v2 · Top 6%#3 Q1 Bottom 39%#39 of 62 95.1% Reasoning effort: Thinking Q2 80.8% Q3 83.5% Q4 79.8% Q5 90.2% Q6 74.6% About Japanese MT-Bench More about Japanese MT-Bench