Anthropic's Responsible Scaling Policy
Key point
Anthropic has announced a 'Responsible Scaling Policy (RSP)' that introduces AI Safety Levels (ASL) to manage catastrophic risks from AI.
Details
Anthropic has announced a Responsible Scaling Policy (RSP) to manage catastrophic risks that could arise as AI model capabilities improve. This policy focuses on preventing large-scale disasters such as terrorists' use of AI to develop biological weapons or AI going out of control autonomously.
The core framework, AI Safety Levels (ASL), was designed based on the Biosafety Level (BSL) standard used for managing biological risks. It requires differentiated safety and security standards depending on a model's potential level of risk.
The ASL level system is as follows:
- ASL-1: Systems with no significant catastrophic risk (e.g., 2018-era LLMs, chess AI)
- ASL-2: Early signs of dangerous capabilities appear, but they are not yet practical (most current LLMs, including Claude, fall into this category)
- ASL-3: Risk significantly increases compared to non-AI-based tools, or the system shows a low level of autonomy. Extremely strict security and verification through red-team testing are essential.
- ASL-4 and above: Not yet defined, but refers to a stage where the potential for catastrophic misuse and autonomy increase qualitatively and sharply.
Anthropic may pause model training if safety measures fail to keep pace with the speed of scaling. This is aimed at balancing the advancement of safety technology with improvements in model performance, by ensuring that safety issues must be resolved before more powerful models can be developed.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.