Estimating worst case frontier risks of open weight LLMs
OpenAI Blog
Researchers tested open-weight large language models by fine-tuning them to maximize capabilities in biology and cybersecurity to identify potential worst-case risks. They evaluated performance across both domains to establish a frontier of what malicious actors could achieve with publicly available model weights. The findings inform decisions about which models should remain restricted versus released as open weights.
Why it matters
In this paper, we study the worst-case frontier risks of releasing gpt-oss. We introduce malicious fine-tuning (MFT), where we attempt to elicit maximum capabilities by fine-tuning gpt-oss to be as capable as possible in two domains: biology and cybersecurity.