AI Briefing
KO

gpt-oss-safeguard Technical Report

·2025.10.29 09:00

Key point

OpenAI has released gpt-oss-safeguard, an open-weight reasoning model designed for policy-based content classification.

Details

gpt-oss-safeguard-120b and gpt-oss-safeguard-20b are open-weight reasoning models post-trained on top of the gpt-oss model. These models are designed to classify content according to policies provided by the user.

Released under the Apache 2.0 license, these models have the following key features:

  • Text-only models compatible with the Responses API
  • Full Chain-of-Thought (CoT) provided and customizable
  • Support for various reasoning efforts (low, medium, high) and support for Structured Outputs

OpenAI recommends using these models for policy-based content classification rather than as core functionality for general user interactions. In addition, the models' capabilities were validated through initial evaluation results on safety and multilingual performance in chat settings.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.