AI Briefing
KO

Strengthening AI Safeguards Through Collaboration with US CAISI and UK AISI

·2025.09.13 04:19

Key point

Anthropic worked with US and UK AI security institutions to test vulnerabilities in Claude models and strengthen safeguards.

Details

Over the past year, Anthropic has collaborated with the US CAISI (Center for AI Standards and Innovation) and the UK AISI (AI Security Institute) to measure and improve the security of AI systems. This partnership, which began at an initial advisory stage, evolved into a process where government teams gain direct access to systems at various stages of model development to conduct ongoing testing.

Government agencies combine their expertise in cybersecurity, intelligence analysis, and threat modeling with machine learning techniques to evaluate sophisticated attack vectors and defense mechanisms. This collaboration focused on identifying vulnerabilities by applying Constitutional Classifiers, which prevent jailbreaks, to the Claude Opus 4 and 4.1 models.

Key identified vulnerabilities and improvements include:

  • Prompt Injection: Discovered and patched a vulnerability that used specific annotations to evade classifier detection.
  • Stress-testing the safeguard architecture: Discovered a sophisticated universal jailbreak that evaded standard detection methods, leading to a fundamental restructuring of the architecture itself, beyond individual patches.
  • Cipher-based attacks: Improved the system to detect and block attempts to bypass safeguards through encryption or character substitution.
  • Input/output obfuscation attacks: Identified attacks that fragment harmful strings into harmless components, strengthening the filtering mechanism.
  • Automated attack refinement: Leveraged jailbreak techniques generated by automated systems that optimize attack strategies to improve defenses.

This collaboration went beyond simply discovering vulnerabilities, contributing to strengthening the overall security approach, including evidence requirements, deployment monitoring, and rapid response capabilities.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.