Arena Software & IT Services
In plain words · Which AI people preferred on expert-level software and IT services questions
- What does it measure?
- Which AI's answers were preferred on expert-level questions in Software & IT Services. Examples: system design, code review, incident root-cause analysis. These are questions the people who actually do that work would ask, not general ones.
- Who evaluates it, and how?
- On Arena Intelligence, a user enters a question, two anonymous AIs answer, and the user votes for the better one. Over 82 million votes are aggregated with a statistical model (Bradley-Terry) into an Elo-style score, and this table takes the 'style-controlled' version from the official dataset, which statistically removes the bias toward long, nicely formatted answers. An AI classifier judges each prompt's reasoning depth and expertise and tags only about 5.5% as 'expert' prompts, which are then divided into occupational categories based on the US Bureau of Labor Statistics classification (SOC), launched November 2025. This table counts only prompts classified as 'Software & IT Services'.
- Reading the score
- This is a head-to-head record, not an exam score (accuracy %). It is computed from the probability of being preferred over other AIs, so there is no maximum, and when a new AI enters, existing scores shift slightly. Typical values are 1,000–1,500; a 100-point gap means the leader is preferred about 64% of the time, a 200-point gap about 76%. Correctness is not counted — a wrong but friendly answer can win — so if accuracy matters, also look at accuracy tests (GPQA, AIME, etc.). Gaps of a few dozen points are effectively the same tier even if the rank differs. This category has few prompts, so top ranks change often; look at score bands rather than ranks.
Last updated: 2026-10-02
Rank
Top score1563Gemini 4 Argon
Models tested97
Last updated2026.10.02
Model release date (newest on the right)
Best score so farHigher on the chart is better
Can it operate a computer itself to finish hours-long development jobs?Terminal-Bench 4.0
Claude Sonnet 5.5🥇
Claude Sonnet 5.5🥇Which AI gives the coding answers people prefer most?Arena Coding
Gemini 4 Argon🥇
Gemini 4 Argon🥇Can it run and fix code until the task is done?LiveBench Agentic Coding
DeepSeek V4.1 Flash🥇
DeepSeek V4.1 Flash🥇
- 🥇GoogleGemini 4 Argon1563
- 🥈AnthropicClaude Fable 51539
- 🥉OpenAIGPT-6.1 Sol1536
- #4AnthropicClaude Opus 4.61535
- #5AnthropicClaude Opus 4.71531
- #6MetaMuse Spark 1.31531
- #7AnthropicClaude Opus 5.51530
- #8XiaomiMiMo-V2.6-Pro1529
- #9AnthropicClaude Fable 5.11528
- #10MetaMuse Spark 1.21528
- #11Moonshot AIKimi K31526
- #12MetaMuse Spark 1.11525
- #13OpenAIGPT-6 Astra1523
- #14MetaMuse Spark1523
- #15OpenAIGPT-5.6 Sol1522
- #16AnthropicClaude Sonnet 5.51522
- #17AnthropicClaude Opus 4.81518
- #18DeepSeekDeepSeek V4.1 Flash1517
- #19GoogleGemini 3.8 Flash1517
- #20AnthropicClaude Opus 51516
- #21AnthropicClaude Sonnet 4.61514
- #22Z.aiGLM 5.3 Flash1514
- #23GoogleGemini 3.1 Pro1513
- #24Z.aiGLM 5.31512
- #25GoogleGemini 3.7 Flash1512
- #26GoogleGemini 3.6 Flash1511
- #27AlibabaQwen3.7 Max1510
- #28XiaomiMiMo V2.5 Pro1509
- #29XiaomiMiMo-V2.6-Flash1508
- #30AnthropicClaude Opus 4.51508
- #31OpenAIGPT-5.41507
- #32Z.aiGLM 5.21506
- #33OpenAIGPT-5.51506
- #34OpenAIGPT-6 Sol1505
- #35OpenAIGPT-5.6 Terra1505
- #36xAIGrok 4.201504
- #37AnthropicClaude Sonnet 51503
- #37Moonshot AIKimi K2.61503
- #39ByteDanceDola Seed 2.0 Pro1503
- #39xAIGrok 4.20 (Reasoning)1503
- #39xAIGrok 4.51503
- #42Z.aiGLM-5.11502
- #43GoogleGemini 3 Flash1500
- #44AlibabaQwen3.6 Max1498
- #45AnthropicClaude Sonnet 4.51497
- #46GoogleGemini 3.5 Flash1497
- #47OpenAIGPT-6 Luna1494
- #48OpenAIGPT-5.6 Luna1493
- #49xAIGrok 4.61493
- #50GoogleGemma 4 31B1493
- #51AnthropicClaude Opus 4.11492
- #52DeepSeekDeepSeek V4 Pro1492
- #53XiaomiMiMo V2 Pro1491
- #54AlibabaQwen3.7 Plus1490
- #55StepfunStep-51490
- #56Moonshot AIKimi K2.51489
- #57MeituanLongcat Flash Chat1489
- #58GoogleGemini 3.5 Flash-Lite1488
- #59Z.aiGLM-51488
- #60OpenAIGPT-5.4 Mini1485
- #61MiniMaxMiniMax M31485
- #62BaiduERNIE 5.0 Thinking1484
- #63AlibabaQwen3.6 Plus1483
- #64AlibabaQwen3.5 397B A17B1481
- #65xAIGrok 4.31479
- #65XiaomiMiMo V2.51479
- #67xAIGrok 4.71476
- #68Z.aiGLM 5V Turbo1473
- #69DeepSeekDeepSeek V4 Flash1471
- #70Mistral AIMistral Medium 3.51469
- #71MiniMaxMiniMax M2.71467
- #72AnthropicClaude Opus 41467
- #73AnthropicClaude Haiku 4.51464
- #74NVIDIANemotron 3 Ultra1463
- #75GoogleGemini 2.5 Pro1460
- #76DeepSeekDeepSeek V3.21458
- #77xAIGrok 4.1 Fast (Reasoning)1457
- #78GoogleGemini 3.1 Flash Lite1455
- #79OpenAIGPT-51454
- #80OpenAIGPT-5.4 Nano1450
- #81TencentHy31446
- #81UpstageSolar Pro 41446
- #83AnthropicClaude Sonnet 41445
- #84MiniMaxMiniMax M2.51437
- #85GoogleGemini 2.5 Flash1421
- #86OpenAIGPT-5 Mini1416
- #87Arcee AITrinity Large Thinking1405
- #88NVIDIANemotron 3 Super1404
- #89GoogleGemini 2.5 Flash Lite1399
- #90AmazonNova 2 Lite1386
- #91OpenAIGPT OSS 120B1385
- #92OpenAIGPT-5 Nano1372
- #93MetaLlama 4 Maverick1358
- #94NVIDIANemotron 3 Nano 30B A3B1355
- #95IBMGranite 4.1 8B1352
- #96MetaLlama 4 Scout1349
- #97OpenAIGPT-4.11324
Half of models ≤ 1493Measured by: User preference votesLast updated 2026-10-02
18 of the 37 AIs released in the last 60 days have no score here yet · Claude Haiku 5.5, Decider V1.1 27B, GPT-6 Luna Decisions and 15 more
Source: Arena Intelligence Data sources & removal requestsWhat to keep in mind
- It’s a preference vote, not a correctness check. A wrong but friendly answer can win.
- Which questions got asked depends on who voted — they may differ from your use.
- Scores labelled 'Self-reported' were published by the AI maker itself, not measured by a third party.