AI Briefing
KO

Anthropic CVP Releases 6 Rounds of Evaluations

·2026.04.29 10:59

Key point

Anthropic CVP released 6 rounds of evaluations of 4 Claude models.

Details

Sunglasses released 6 evaluation rounds over the past 10 days, repeatedly applying the same prompt set to 4 Claude models.

  • Target models: Opus 4.7, Opus 4.6, Sonnet 4.6, Haiku 4.5
  • Results: 120 transcripts were produced across 10 model-effort combinations, and all 120 were rated clean.
  • Method: 3 fixed prompts were used before execution, run in isolated Claude Code sessions, with transcripts captured, hashed, and archived, and responses graded on allowed / partial / blocked, usefulness, and safety criteria.
  • Timeline: approved on April 16, 2026, the 6th run concluded on April 26, and the family synthesis report was released on April 27
  • Next steps: an adversarial-framing appendix probe set will be used to conduct additional checks with JailbreakBench, HarmBench, AdvBench, PromptInject, Garak, PyRIT, and recent CVE PoCs.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.