AI Briefing
KO

Open-source AI alignment tool handed off

·2026.05.07 00:00

Key point

Anthropic released Petri 3.0 and transferred its development to Meridian Labs.

Details

Anthropic has used Petri, an open-source alignment testing tool released in October 2025, to quickly check any LLM for tendencies toward deception, sycophancy, and compliance with harmful requests. Petri works by having a separate auditor model create scenarios, observing the target model's responses, and then having a judge model score the resulting conversations.

Anthropic has included Petri in the alignment evaluation of all Claude models since Claude Sonnet 4.5, and the UK's AI Security Institute (AISI) has also used Petri as a core tool for assessing models' potential to obstruct research.

The changes in this Petri 3.0 release are as follows.

  • Adaptability: The auditor model and target model have been separated, allowing each component to be adjusted independently.
  • Realism: An additional tool called Dish uses real system prompts and real scaffolds to increase test realism.
  • Depth: Integration with Bloom enables deeper in-depth analysis.

Anthropic has handed off Petri's development to Meridian Labs, a nonprofit AI evaluation organization. Petri will now be maintained as an independent AI evaluation tool separate from any specific AI lab, forming, together with Inspect and Scout, a shared evaluation stack that researchers, governments, and AI labs can all use.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.