AI Briefing
KO

Testing Safety Defenses Through a New Bug Bounty Program

·2025.07.24 03:02

Key point

Anthropic is introducing a new bug bounty program and recruiting security researchers to strengthen AI safety.

Details

Anthropic has launched a new bug bounty program to stress-test its latest safety measures. This program aims to find universal jailbreak cases in as-yet-unreleased Safety Classifiers.

This is part of the Responsible Scaling Policy, a process for meeting AI Safety Level 3 (ASL-3) deployment standards. Run in partnership with HackerOne, the program specifically tests an updated version of the Constitutional Classifiers system, which prevents the leakage of CBRN (chemical, biological, radiological, and nuclear) related information.

Participants receive the following benefits:

  • Early access to Claude 3.7 Sonnet
  • A bounty of up to $25,000 for verified universal jailbreak discoveries

The initial program has now ended, and a new bug bounty program is currently running to test the Claude Opus 4 model and other safety systems. Anthropic also continues to accept reports of ASL-3 related risks (such as biological threats) discovered on public platforms like social media.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.