TLDRocket
Sign in

gpt-oss-safeguard technical report

OpenAI Blog Covered by 2 sources

Two open-weight reasoning models, gpt-oss-safeguard-120b and gpt-oss-safeguard-20b, were developed by post-training gpt-oss models to apply provided safety policies for content labeling. The models come in 120 billion and 20 billion parameter sizes. The technical report establishes baseline safety evaluations comparing the new safeguard models against their underlying gpt-oss predecessors.

Why it matters

gpt-oss-safeguard-120b and gpt-oss-safeguard-20b are two open-weight reasoning models post-trained from the gpt-oss models and trained to reason from a provided policy in order to label content under that policy. In this report, we describe gpt-oss-safeguard’s capabilities and provide our baseline safety evaluations on the gpt-oss-safeguard models, using the underlying gpt-oss models as a baseline. For more information about the development and architecture of the underlying gpt-oss models, see the original gpt-oss model model card⁠.

Also covered by

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads 60+ sources, removes duplicate coverage, and summarises the day in two minutes. Free, no spam, unsubscribe anytime.