AI Briefing
KO

Model Safety Bug Bounty

·2024.09.11 02:28

Key point

Anthropic is launching a new bug bounty program to find 'universal jailbreak' vulnerabilities in its next-generation AI safety systems.

Details

Anthropic is expanding its Bug Bounty program to find flaws in next-generation safety systems designed to prevent AI model misuse. This program focuses on identifying and mitigating Universal Jailbreak attacks that can consistently bypass AI safety guardrails across various domains.

In particular, it aims to address vulnerabilities in high-risk areas such as CBRN (Chemical, Biological, Radiological, and Nuclear) and cybersecurity. Participants will be granted Early Access to next-generation safety mitigation systems that have not yet been publicly released.

Key details of the program are as follows:

  • Reward scale: Up to $15,000 paid for discovering novel, universal jailbreak attacks
  • Operating method: Run in partnership with HackerOne, initially operating on an invite-only basis
  • How to apply: Experts such as AI security researchers can apply via the application form by August 16

This initiative aligns with international efforts toward responsible AI development, including the White House's Voluntary AI Commitments and the G7 Hiroshima Process's Code of Conduct for Organizations Developing Advanced AI Systems.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.