Strengthening the Safety Ecosystem Through External Testing
Key point
OpenAI is strengthening an independent evaluation system through collaboration with external expert organizations to verify the safety of frontier AI.
Details
OpenAI actively leverages independent and trustworthy Third-party assessments to strengthen the safety ecosystem for frontier AI. The purpose is to verify claims about a model's safety capabilities and mitigation measures, prevent potential blind spots, and increase transparency.
External collaboration takes three main forms.
- Independent evaluations: Testing core risk areas such as biosecurity, cybersecurity, AI self-improvement, and scheming
- Methodology reviews: Verifying risk assessment and interpretation methods
- SME probing: Domain experts directly evaluate models through real-world tasks and provide feedback on safety measures
These evaluations serve as an independent layer that complements the limitations of internal testing and prevents Self-confirmation bias. For GPT-5, a wide range of external capability evaluations were coordinated to support deployment decisions, covering areas such as long-horizon autonomy, deception, oversight subversion, and potential for biological experiment planning.
The evaluation process utilizes benchmarks such as METR's time horizon evaluations and SecureBio's Virology Capabilities Test (VCT). OpenAI supports external organizations by providing secure access to early model checkpoints, a Zero-data retention environment, and models without safety mitigations applied, enabling precise measurement of a model's underlying capabilities.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.