Close Menu
AI News TodayAI News Today

    Subscribe to Updates

    Get the latest creative news from FooBar about art, design and business.

    What's Hot

    OpenAI agent “didn’t accept no for an answer” in Australian government breach

    Google’s first Suncatcher orbital data center test launches October 1

    How to Use PopAI to Turn a Messy Idea into a Polished Deck in 45 Minutes

    Facebook X (Twitter) Instagram
    • About Us
    • Contact Us
    Facebook X (Twitter) Instagram Pinterest Vimeo
    AI News TodayAI News Today
    • Home
    • AI News
    • AI Reviews
    • AI Tools
    • AI Tutorials
    • Chatbots
    • Free AI Tools
    • Artificial Intelligence
    AI News TodayAI News Today
    Home»AI Tutorials»How UK AISI and EvalEval Are Making Benchmark Results Reproducible
    AI Tutorials

    How UK AISI and EvalEval Are Making Benchmark Results Reproducible

    By No Comments4 Mins Read
    Share Facebook Twitter Pinterest LinkedIn Tumblr Reddit Telegram Email
    How UK AISI and EvalEval Are Making Benchmark Results Reproducible
    Share
    Facebook Twitter LinkedIn Pinterest Email


    The EvalEval Coalition is thrilled to share that the UK AI Security Institute (AISI) is using EvalEval’s infrastructure to openly share evaluation results, supporting more reproducible and verifiable evaluation science.

    AISI and EvalEval have previously collaborated on research that began at a joint workshop alongside NeurIPS 2025, and feedback from the Institute has helped shape the Every Eval Ever (EEE) schema. This next phase of the collaboration puts that shared infrastructure into practice.



    Why reproducible evaluation reporting matters

    As AI deployment accelerates, evaluations are becoming increasingly important sources of evidence about model and system performance. Yet results are reported across many formats, platforms, and outlets, often without enough information to reproduce them. Running the evaluations again may itself be prohibitively expensive.

    EvalEval’s mission is to improve this ecosystem through a shared reporting schema, Every Eval Ever, and an open platform, Evaluation Cards, that brings evaluation results and the information needed to interpret them into a common structure.

    This builds naturally on AISI’s work to make evaluation more efficient through OptStop, more statistically rigorous through HiBayES, and more standardised in areas including transcript analysis and capability elicitation. Together, AISI and EvalEval are working to diagnose gaps in evaluation reporting and build shared infrastructure to close them.



    What AISI is sharing

    Transcript-level transparency matters not only for reproducibility, but also for analysis and diagnosis. In this new phase of the collaboration, AISI is making publicly reported evaluation methods and findings available through Evaluation Cards where appropriate. The release includes verified results, context, and configuration information for the five benchmarks in the paper’s main experiment:

    • HealthBench
    • FrontierMath
    • Humanity’s Last Exam
    • SWE-Bench Pro
    • Terminal-Bench 2.0

    These results cover six frontier models: Claude Opus 4, Claude Opus 4.5, Claude Opus 4.6, GPT-5, GPT-5.2, and GPT-5.4. The release also includes results from two related cyber evaluations—Cyber CTFs and The Last Ones—which use a different, partially overlapping set of models. The data accompany AISI’s paper, How Inference Compute Shapes Frontier LLM Evaluation, which studies how benchmark performance depends on inference-time compute and evaluation protocol.

    Performance on Humanity’s Last Exam changes with evaluation protocol and inference compute. Each curve shows the cumulative share of attempted tasks solved within a given token count, using the earliest observed success per task. When models received correctness feedback from an oracle after each attempt, they continued to solve additional tasks as token use increased.

    When results are openly released with setup information, researchers and practitioners can examine individual studies more closely and compare findings across the wider ecosystem. Where other reports lack these details, releases like AISI’s provide verified reference points for interpreting evaluations in context—for example, by helping researchers understand how setup choices may influence reported performance. As more evaluators adopt EEE, open comparisons like these can support broader and more reliable meta-research.

    AISI’s Terminal-Bench 2.0 results alongside other reported evaluations for the same models, under different evaluation setups.

    We are excited about this adoption and look forward to further standardising and sharing evaluations with AISI and other AI evaluation organisations.



    Contribute to the shared mission



    About the EvalEval Coalition

    The EvalEval Coalition is a research community developing scientifically grounded research and robust deployment infrastructure for the evaluation ecosystem. Its goal is to improve evaluation science, address the lack of consensus around documenting evaluation applicability and utility, and broaden coverage of the impacts that matter for scientific research and policy analysis.

    The coalition’s flagship projects include Every Eval Ever, a shared schema and repository for evaluation results, and Evaluation Cards, which combines benchmark metadata, evaluation-run data, and model metadata into interpretable records. Together, they make it easier to understand when apparently similar scores were produced under meaningfully different conditions.



    About the UK AI Security Institute

    The UK AI Security Institute is a research organisation within the UK government’s Department for Science, Innovation and Technology. Its mission is to equip governments with a scientific understanding of the risks posed by advanced AI. AISI conducts research and builds infrastructure to understand advanced AI capabilities and impacts, develop and test mitigations, and inform policy.



    Further reading

    AISI benchmark EvalEval making Reproducible results
    Share. Facebook Twitter Pinterest LinkedIn Tumblr Email
    Previous ArticlePitPro’s first tire-changing robot goes live in Canada
    Next Article From Words to Vectors: What Happens in Between?
    • Website

    Related Posts

    AI Tutorials

    Stanford CS231n: What You Actually Learn (And Whether It’s Still Worth It)

    AI Tutorials

    Stanford CS229: The Machine Learning Course That Changed Everything

    Chatbots

    Meta is making a standalone Muse AI gadget

    Add A Comment
    Leave A Reply Cancel Reply

    Top Posts

    OpenAI agent “didn’t accept no for an answer” in Australian government breach

    0 Views

    Google’s first Suncatcher orbital data center test launches October 1

    0 Views

    How to Use PopAI to Turn a Messy Idea into a Polished Deck in 45 Minutes

    0 Views
    Stay In Touch
    • Facebook
    • YouTube
    • TikTok
    • WhatsApp
    • Twitter
    • Instagram
    Latest Reviews
    AI Tutorials

    Quantization from the ground up

    AI Tools

    David Sacks is done as AI czar — here’s what he’s doing instead

    AI Reviews

    Judge sides with Anthropic to temporarily block the Pentagon’s ban

    Subscribe to Updates

    Get the latest tech news from FooBar about tech, design and biz.

    Most Popular

    OpenAI agent “didn’t accept no for an answer” in Australian government breach

    0 Views

    Google’s first Suncatcher orbital data center test launches October 1

    0 Views

    How to Use PopAI to Turn a Messy Idea into a Polished Deck in 45 Minutes

    0 Views
    Our Picks

    Quantization from the ground up

    David Sacks is done as AI czar — here’s what he’s doing instead

    Judge sides with Anthropic to temporarily block the Pentagon’s ban

    Subscribe to Updates

    Get the latest creative news from FooBar about art, design and business.

    Facebook X (Twitter) Instagram Pinterest
    • About Us
    • Contact Us
    • Terms & Conditions
    • Privacy Policy
    • Disclaimer

    © 2026 ainewstoday.co. All rights reserved. Designed by DD.

    Type above and press Enter to search. Press Esc to cancel.