AISI Publishes Reproducible AI Evaluations
- •UK AI Security Institute uses EvalEval infrastructure to publish reproducible AI evaluation results
- •Release covers 5 benchmarks, 6 frontier models and 2 related cyber evaluations
- •AISI paper studies how inference compute and evaluation protocol shape frontier LLM results
The UK AI Security Institute is using EvalEval infrastructure to publish evaluation results in a more reproducible and verifiable format, according to a September 22, 2026 Hugging Face blog post by members of EvalEval and AISI. The collaboration moves earlier joint work into practice after research that began at a joint workshop alongside NeurIPS 2025, where AISI feedback helped shape the Every Eval Ever schema. EvalEval says the goal is to put evaluation results, context and interpretation details into a shared structure through Every Eval Ever and Evaluation Cards.
EvalEval argues that reproducible reporting matters because AI evaluation results now serve as evidence about model and system performance, but are often spread across many formats, platforms and outlets without enough detail to reproduce them. Re-running evaluations can also be prohibitively expensive. The group says shared reporting can help researchers understand how evaluation setup choices affect reported scores, especially when other reports do not include comparable configuration details.
AISI is making publicly reported methods and findings available through Evaluation Cards where appropriate. The release includes verified results, context and configuration information for five benchmarks in the paper's main experiment: HealthBench, FrontierMath, Humanity's Last Exam, SWE-Bench Pro and Terminal-Bench 2.0. The results cover six frontier models: Claude Opus 4, Claude Opus 4.5, Claude Opus 4.6, GPT-5, GPT-5.2 and GPT-5.4. AISI also released results from two related cyber evaluations, Cyber CTFs and The Last Ones, using a different, partially overlapping set of models.
The data accompany AISI's paper, "How Inference Compute Shapes Frontier LLM Evaluation," which studies how benchmark performance changes with inference-time compute and evaluation protocol. The blog says Humanity's Last Exam performance varies by evaluation setup and token count; when models received correctness feedback from an oracle after each attempt, they kept solving additional tasks as token use increased. AISI's Terminal-Bench 2.0 results are also shown beside other reported evaluations for the same models under different setups.
EvalEval says broader adoption of Every Eval Ever could support more reliable meta-research by enabling open comparisons across studies. The coalition asks model developers to report verified evaluation results, evaluation developers to report benchmarks and run data using the Every Eval Ever schema, and evaluation, governance and policy researchers to explore Evaluation Cards by benchmark or model. EvalEval describes itself as a research community building evaluation infrastructure, while AISI is a UK government research organisation within the Department for Science, Innovation and Technology focused on risks from advanced AI.