Jev Trails Opus on Classification Accuracy
- •Jev 1.13 scored 81.0% accuracy versus Claude Opus 5 at 84.4% on Banking77 classification
- •Jev ran 13 times faster at median latency and cost $0.11 per 1,000 requests
- •A 0.90 confidence cascade reached 84.0% accuracy at $0.69 per 1,000 requests
OpenRouter tested TypeSafe’s Jev 1.13 against Claude Opus 5 on 22 September 2026 using 3,080 Banking77 customer support utterances, with each utterance classified into one of 77 banking intents. Jev is TypeSafe’s System One decision model: it takes an app state object and typed question, then returns a typed answer, confidence score, and probability for each option through the Decisions API. Claude Opus 5 was chosen because OpenRouter users spent the most on it for classification in the Task spend section of OpenRouter’s rankings page as of 22 September 2026.
Jev 1.13 scored 81.0% accuracy with a 95% interval of 79.6 to 82.3, while Claude Opus 5 scored 84.4% with a 95% interval of 83.1 to 85.6. Jev recorded 80.5% macro-F1 (equal weight across classes), compared with 83.6% for Claude Opus 5. Both models returned 0 invalid responses out of 3,080, so every error was a wrong label rather than a malformed response.
The speed and cost differences were larger than the accuracy gap. Jev’s latency was 175 ms at p50, 270 ms at p95, and 353 ms at p99, while Claude Opus 5 measured 2,266 ms at p50, 3,004 ms at p95, and 3,835 ms at p99. Jev cost $0.34 for the full run, or $0.11 per 1,000 requests; Claude Opus 5 cost $7.44, or $2.42 per 1,000 requests, with prompt caching enabled on a roughly 3,700-token system prompt.
OpenRouter used the full Banking77 test split from PolyAI, which has 3,080 examples, 40 per intent, under the CC BY 4.0 license. Each utterance was short, with a median length of nine words. The same one-line intent criteria, written only from label names and without inspecting test data, were given to both models; Jev received a Decisions API Choice question, while Claude Opus 5 received a prompt with a strict JSON schema response format, temperature zero, reasoning turned off, and 8 concurrent requests from the same machine.
Claude Opus 5 led Jev by 3.3 percentage points on accuracy, and a paired bootstrap placed the 95% confidence interval for that gap at 2.3 to 4.4 points. The two models agreed on 89.3% of utterances. Where their answers differed, Claude Opus 5 alone was correct on 175 examples, while Jev alone was correct on 72 examples.
Class-level results varied. Claude Opus 5 led Jev on 35 of 77 intents, Jev led on 15, and they tied on 27. Claude Opus 5’s largest advantage was on receiving_money, scoring 80.0% versus Jev’s 52.5%. Jev’s strongest listed advantage was on compromised_card, scoring 95.0% while Claude Opus 5 scored 70.0%, with the article saying Opus often confused stolen card details with an unrecognized payment.
Jev’s confidence score was not calibrated as a probability, but it ranked examples usefully. For the 58% of utterances with confidence ≥0.99, Jev reached 96.3% accuracy; for the 3.5% below 0.5 confidence, accuracy fell to 29.6%. A cascade (routing uncertain cases onward) at threshold 0.90 sent 75.9% of traffic through Jev, reached 84.0% accuracy, and cost $0.69 per 1,000 requests, compared with Claude Opus 5 alone at 84.4% and $2.42 per 1,000 requests.
OpenRouter listed several caveats: the test used one dataset, one domain, one prompt design, and one fifteen-minute window on one afternoon. Banking77 was published in 2020, so the article said Claude Opus 5 could have a memorization edge, though the run alone could not prove it. Both models also scored 0 out of 40 on get_physical_card because the label name did not reveal that it covered questions about separate PIN delivery; excluding that class, Jev reached 82.1% and Claude Opus 5 reached 85.5%.