AI Briefing
KO

New Initiative to Develop Third-Party Model Evaluations

·2024.09.11 02:28

Key point

Anthropic has announced a new initiative to support third-party organizations in developing evaluations to measure advanced AI models' capabilities and risks.

Details

Accurately assessing the capabilities and risks of AI models requires a robust third-party evaluation ecosystem. However, demand for high-quality, safety-relevant evaluation development currently outpaces supply.

Anthropic is introducing a new initiative to support third-party organizations developing evaluation tools that can effectively measure the capabilities of advanced AI models. This investment aims to raise the bar for the AI safety field as a whole and provide tools that benefit the broader ecosystem.

The key priority areas are as follows:

  • AI Safety Level (ASL) evaluations: evaluations to measure the safety levels defined in the Responsible Scaling Policy (cybersecurity, CBRN risks, model autonomy, national security risks, etc.)
  • Development of advanced capability and safety metrics
  • Infrastructure, tools, and methodologies for evaluation development

In particular, in the cybersecurity domain, the focus is on assessing capabilities at the level of sophisticated threat actors, such as vulnerability discovery and exploit development, while in the CBRN (chemical, biological, radiological, nuclear) risk domain, the focus is on measuring models' ability to design lethal threats or amplify the capabilities of non-experts.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.