gpt-oss-safeguard technical report
OpenAI Blog ● Covered by 2 sources
Two open-weight reasoning models, gpt-oss-safeguard-120b and gpt-oss-safeguard-20b, were developed by post-training gpt-oss models to apply provided safety policies for content labeling. The models come in 120 billion and 20 billion parameter sizes. The technical report establishes baseline safety evaluations comparing the new safeguard models against their underlying gpt-oss predecessors.
Why it matters
gpt-oss-safeguard-120b and gpt-oss-safeguard-20b are two open-weight reasoning models post-trained from the gpt-oss models and trained to reason from a provided policy in order to label content under that policy. In this report, we describe gpt-oss-safeguard’s capabilities and provide our baseline safety evaluations on the gpt-oss-safeguard models, using the underlying gpt-oss models as a baseline. For more information about the development and architecture of the underlying gpt-oss models, see the original gpt-oss model model card.