OpenAI's Safety Practices
·2024.05.21 15:00
Key point
At the AI Seoul Summit, OpenAI revealed key safety practices applied from model development through deployment.
Details
Through the AI Seoul Summit, OpenAI shared a scientific and balanced approach to securing both model performance and safety at the same time. The goal is to integrate safety from the earliest stages of model development to achieve both innovation and reliability simultaneously.
The key practices are as follows:
- Empirical Red-Teaming: Under the Preparedness Framework, safety is evaluated by internal and external experts before model release. For GPT-4o, more than 70 external experts assessed risks.
- Alignment and Safety Research: Heavy investment in post-training research to reduce factual errors in models and increase robustness against attacks such as jailbreak.
- Abuse Monitoring: Risks are detected using a dedicated moderation model and GPT-4, and risk mitigation information is shared, including jointly disclosing cases of state-sponsored hacking groups abusing the technology together with Microsoft.
- Systematic Safety Approach: Safety measures are implemented across the entire model lifecycle, from pre-training to deployment.
- Child Protection: Strong guardrails are built into ChatGPT and DALL·E, and OpenAI collaborates with Thorn to detect and report harmful content related to children.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.