Responsible Scaling Policy V3
Key point
Anthropic has released the third version of the 'Responsible Scaling Policy (RSP),' a voluntary framework for mitigating the catastrophic risks of AI.
Details
Anthropic has released the third version of the Responsible Scaling Policy (RSP), a voluntary framework for mitigating the catastrophic risks of AI systems. This update draws on two years of operational experience to reinforce the strengths of the existing policy while including new measures to increase transparency and accountability in decision-making.
The core principle of the RSP is an if-then commitment. If an AI model exceeds a certain capability level (for example, the ability to assist in creating biological weapons), it is designed to introduce stricter safeguards to prevent model misuse or weight theft.
Safety measures are applied in stages according to the AI Safety Level (ASL). Detailed guidelines are currently defined for ASL-2 and ASL-3, while ASL-4 and beyond, intended for next-generation models, will be specified in line with technological progress.
Through the RSP, Anthropic aims to drive the following changes:
- Internal enforcement capability: Making safeguards a mandatory requirement during the training and release process of new models.
- Race to the top: Encouraging the industry as a whole to raise safety standards, forming voluntary industry norms.
- Building consensus on risk: When specific capability thresholds are reached, using this as a basis to pursue multilateral cooperation with governments and other companies.
- Preparing for the future: Cooperating on global-level safety measures in preparation for high-risk situations that individual companies would find difficult to address alone.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.