AI Briefing
KO

AISI and EvalEval Collaborate to Standardize LLM Evaluation Reproducibility

·2026.09.22 09:00

Key point

The UK's AISI and EvalEval have established a collaborative framework to enhance the reproducibility and verifiability of LLM evaluation results through the Every Eval Ever schema.

1 / 2

Details

The UK's AI Security Institute (AISI) and the EvalEval Coalition have collaborated to build infrastructure that improves the reproducibility and verifiability of LLM evaluation results. This collaboration is based on the NeurIPS 2025 workshop and shaped the Every Eval Ever (EEE) schema by incorporating feedback from AISI.

Release of Standardized Evaluation Data

AISI releases key experimental data through Evaluation Cards, which serve as reference points for researchers to deeply analyze individual studies and compare results across the ecosystem. The released data includes:

  • 5 Major Benchmarks: HealthBench, FrontierMath, Humanity's Last Exam, SWE-Bench Pro, Terminal-Bench 2.0
  • 6 Target Models: Claude Opus 4, 4.5, 4.6, GPT-5, 5.2, 5.4
  • 2 Cyber Evaluations: Cyber CTFs, The Last Ones
  • Data Source: AISI paper 'How Inference Compute Shapes Frontier LLM Evaluation' (research on the impact of inference compute resources and evaluation protocols on performance)

Emphasis on the Importance of Evaluation Protocols

The released data demonstrates that performance can vary significantly depending on evaluation settings. For example, in Humanity's Last Exam, it was found that receiving correct answer feedback from an Oracle allows models to solve additional problems as token usage increases. Additionally, results for Terminal-Bench 2.0 are provided in a way that allows comparison with results from other evaluation settings, helping to understand the impact of setting choices on performance reporting.

Goals of the EvalEval Coalition

The EvalEval Coalition is a community developing scientific research and robust deployment infrastructure for the evaluation ecosystem. Its main goals are improving evaluation science, addressing the lack of consensus on documenting evaluation applicability, and expanding the scope of impact related to policy analysis. Every Eval Ever provides a shared schema and repository for evaluation results, while Evaluation Cards combine benchmark metadata, evaluation run data, and model metadata to clearly distinguish similar scores under different conditions.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.