AI Briefing
KO

OpenAI and Anthropic Share Joint Safety Evaluation Results

·2025.08.27 19:00

Key point

OpenAI and Anthropic have released the results of a joint safety and alignment evaluation of each other's models.

Details

OpenAI and Anthropic conducted their first-ever joint evaluation, using each company's internal safety and misalignment evaluation tools to test the other's publicly released models. The collaboration was pursued to identify potential flaws in models early through transparent evaluation and to advance safety and alignment techniques.

The models evaluated included Claude Opus 4, Claude Sonnet 4, and GPT-4o, GPT-4.1, OpenAI o3, OpenAI o4-mini. To improve the accuracy of testing, both companies conducted the evaluation with some of the models' external safeguards relaxed, and also analyzed differences depending on whether the models' reasoning functionality was enabled.

The key evaluation results are as follows.

  • Instruction Hierarchy: Claude 4 models showed excellent performance in preventing conflicts between system messages and user messages, and also demonstrated strong results in resistance to system prompt extraction.
  • Jailbreaking: Claude models showed relatively weaker defensive performance compared to OpenAI o3 and o4-mini. In certain scenarios, Claude models with reasoning disabled actually performed better.
  • Hallucination: Claude models recorded very high refusal rates, reaching up to 70%, during the hallucination evaluation.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.