Compare AIFind AIAI NewsAI How-To
About Us
PrivacyTermsFAQContactContact
AIB Inc.Company info
© 2026 AIB Inc.

Harvey Updates Legal AI Benchmark Scoring

Harvey Updates Legal AI Benchmark Scoring

Artificial Analysis·Friday, October 9, 2026
  • •Harvey and Artificial Analysis update legal-agent benchmark scoring with hallucination audits and a three-judge panel
  • •Grok 4.7 leads the 120-task evaluation at 9.4%, ahead of Muse Spark 1.3 and GPT-6 Astra
  • •More than 60% of otherwise passing results contain material hallucinations across tested models
  • •Harvey and Artificial Analysis update legal-agent benchmark scoring with hallucination audits and a three-judge panel
  • •Grok 4.7 leads the 120-task evaluation at 9.4%, ahead of Muse Spark 1.3 and GPT-6 Astra
  • •More than 60% of otherwise passing results contain material hallucinations across tested models
  • •Harvey and Artificial Analysis update legal-agent benchmark scoring with hallucination audits and a three-judge panel
  • •Grok 4.7 leads the 120-task evaluation at 9.4%, ahead of Muse Spark 1.3 and GPT-6 Astra
  • •More than 60% of otherwise passing results contain material hallucinations across tested models
  • •Harvey and Artificial Analysis update legal-agent benchmark scoring with hallucination audits and a three-judge panel
  • •Grok 4.7 leads the 120-task evaluation at 9.4%, ahead of Muse Spark 1.3 and GPT-6 Astra
  • •More than 60% of otherwise passing results contain material hallucinations across tested models

Artificial Analysis and Harvey released Harvey LAB-AA v1.1 on October 8, 2026, updating the scoring method for an independent benchmark of AI agents completing real-world legal tasks. Each model completes 120 private tasks spanning corporate M&A, capital markets, tax, litigation and bankruptcy. Every deliverable is checked against task source documents for hallucinations, and three LLM judges grade each rubric criterion. A task earns credit in the new Hallucination-Gated All-Pass Rate only if it passes every criterion and contains no material hallucination.

Grok 4.7 (xhigh) led the new metric at 9.4%, followed by Muse Spark 1.3 (max) at 8.9% and GPT-6 Astra (max) at 8.6%. GPT-6.1 Sol (max) scored 6.9%, Claude Fable 5.1 (max, with fallback) 6.4%, Kimi K3 (max) 5.3%, and Claude Opus 5.5 (max, with fallback) 4.2%; the Claude models did not use the enabled fallback. Without the hallucination gate, Muse Spark would lead with a 26.7% All-Pass Rate, but two thirds of those passing results contain a material hallucination. GPT-6 Astra retains nearly all its passes, dropping from 8.9% to 8.6%, and moves from joint 10th place to 3rd; GPT-6.1 Sol falls from 7.5% to 6.9%. Across tested models, more than 60% of otherwise passing results contain a material hallucination, and several score 0% after the gate.

Sixteen models passed 85.6–96.0% of rubric criteria before the hallucination gate. The benchmark reports that Criterion Pass Rate separately from the mean material hallucinations per checked task, to show rubric coverage and source grounding as distinct measures. GPT-6 Astra averaged 0.03 material hallucinations per task, or 4 across all 120 tasks; GPT-6 Sol averaged 0.07, or 8 across 120. Gemini 3.8 Flash averaged the most in the launch set, at 13.96 per task. Muse Spark passed the most criteria, 96.0%, but averaged 1.68 material hallucinations per task, compared with 0.03 for GPT-6 Astra. Every open-weights model averaged at least 2.09 material hallucinations per task. The evaluation says hallucination rates depend more on the model than on the legal practice area.

Version 1.1 replaces the single judge used in v1.0 with a panel of GPT-6 Sol, Grok 4.7 and Claude Opus 5.5; verdicts are averaged across the three. The hallucination check uses GPT-6 Sol (high) in two passes: it flags possible contradictions, fabricated source content or specific claims unsupported by the sources, then rechecks each flag against the documents and classifies upheld errors as material or minor. Only a material hallucination zeros a task’s score; tasks without a usable submission score zero and are not audited. Version 1.1 results are not directly comparable with v1.0 because both scoring and grading changed. The dataset was updated to Harvey’s private v1.1.0 set. Future updates are planned to account for lawyer-valued usability factors such as style and tone.

The independent implementation runs models on Artificial Analysis’s Stirrup agent harness, which supports context compaction, and uses simplified prompts written by Artificial Analysis. It omits Harvey’s custom tools and document-generation scripts, instead giving models a simple code execution tool; deliverables must use the exact requested filename. In a checker comparison on 20 tasks and eight models, GPT-6 Sol upheld 470 material hallucinations, versus 219 for Grok, 99 for Claude Opus and 57 for Claude Sonnet. All six checkers found none in GPT-6 Astra’s outputs. Near-pass scoring allows one or two missed criteria while retaining the hallucination gate: GPT-6 Astra scored 20.3% with one miss and 31.7% with two; GPT-6.1 Sol scored 20.3% and 29.6%, while Grok scored 15.3% and 23.1%. Grok cost about $9.50 per task, less than half Claude Fable’s roughly $21.70; Muse Spark was second for about $4.20. GPT-6 Astra used about 81k output tokens per task and scored 8.6%, compared with Grok’s roughly 180k tokens.

Artificial Analysis and Harvey released Harvey LAB-AA v1.1 on October 8, 2026, updating the scoring method for an independent benchmark of AI agents completing real-world legal tasks. Each model completes 120 private tasks spanning corporate M&A, capital markets, tax, litigation and bankruptcy. Every deliverable is checked against task source documents for hallucinations, and three LLM judges grade each rubric criterion. A task earns credit in the new Hallucination-Gated All-Pass Rate only if it passes every criterion and contains no material hallucination.

Grok 4.7 (xhigh) led the new metric at 9.4%, followed by Muse Spark 1.3 (max) at 8.9% and GPT-6 Astra (max) at 8.6%. GPT-6.1 Sol (max) scored 6.9%, Claude Fable 5.1 (max, with fallback) 6.4%, Kimi K3 (max) 5.3%, and Claude Opus 5.5 (max, with fallback) 4.2%; the Claude models did not use the enabled fallback. Without the hallucination gate, Muse Spark would lead with a 26.7% All-Pass Rate, but two thirds of those passing results contain a material hallucination. GPT-6 Astra retains nearly all its passes, dropping from 8.9% to 8.6%, and moves from joint 10th place to 3rd; GPT-6.1 Sol falls from 7.5% to 6.9%. Across tested models, more than 60% of otherwise passing results contain a material hallucination, and several score 0% after the gate.

Sixteen models passed 85.6–96.0% of rubric criteria before the hallucination gate. The benchmark reports that Criterion Pass Rate separately from the mean material hallucinations per checked task, to show rubric coverage and source grounding as distinct measures. GPT-6 Astra averaged 0.03 material hallucinations per task, or 4 across all 120 tasks; GPT-6 Sol averaged 0.07, or 8 across 120. Gemini 3.8 Flash averaged the most in the launch set, at 13.96 per task. Muse Spark passed the most criteria, 96.0%, but averaged 1.68 material hallucinations per task, compared with 0.03 for GPT-6 Astra. Every open-weights model averaged at least 2.09 material hallucinations per task. The evaluation says hallucination rates depend more on the model than on the legal practice area.

Version 1.1 replaces the single judge used in v1.0 with a panel of GPT-6 Sol, Grok 4.7 and Claude Opus 5.5; verdicts are averaged across the three. The hallucination check uses GPT-6 Sol (high) in two passes: it flags possible contradictions, fabricated source content or specific claims unsupported by the sources, then rechecks each flag against the documents and classifies upheld errors as material or minor. Only a material hallucination zeros a task’s score; tasks without a usable submission score zero and are not audited. Version 1.1 results are not directly comparable with v1.0 because both scoring and grading changed. The dataset was updated to Harvey’s private v1.1.0 set. Future updates are planned to account for lawyer-valued usability factors such as style and tone.

The independent implementation runs models on Artificial Analysis’s Stirrup agent harness, which supports context compaction, and uses simplified prompts written by Artificial Analysis. It omits Harvey’s custom tools and document-generation scripts, instead giving models a simple code execution tool; deliverables must use the exact requested filename. In a checker comparison on 20 tasks and eight models, GPT-6 Sol upheld 470 material hallucinations, versus 219 for Grok, 99 for Claude Opus and 57 for Claude Sonnet. All six checkers found none in GPT-6 Astra’s outputs. Near-pass scoring allows one or two missed criteria while retaining the hallucination gate: GPT-6 Astra scored 20.3% with one miss and 31.7% with two; GPT-6.1 Sol scored 20.3% and 29.6%, while Grok scored 15.3% and 23.1%. Grok cost about $9.50 per task, less than half Claude Fable’s roughly $21.70; Muse Spark was second for about $4.20. GPT-6 Astra used about 81k output tokens per task and scored 8.6%, compared with Grok’s roughly 180k tokens.

Read original (English)·Oct 8, 2026
#harvey#lab aa#legal agent benchmark#hallucination gated all pass rate#grok 4 7#gpt 6 astra#muse spark 1 3#legal ai evaluation