TLDRocket
Sign in

Building an early warning system for LLM-aided biological threat creation

OpenAI Blog

Researchers created an evaluation framework to assess whether large language models could help someone develop biological threats, testing GPT-4 with biology experts and students. The testing found GPT-4 provided at most a mild accuracy improvement for biological threat creation tasks, a difference too small to be statistically conclusive. The work establishes a baseline methodology for ongoing assessment of LLM biosecurity risks and calls for further investigation across different models and threat scenarios.

Why it matters

We’re developing a blueprint for evaluating the risk that a large language model (LLM) could aid someone in creating a biological threat. In an evaluation involving both biology experts and students, we found that GPT-4 provides at most a mild uplift in biological threat creation accuracy. While this uplift is not large enough to be conclusive, our finding is a starting point for continued research and community deliberation.

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads 60+ sources, removes duplicate coverage, and summarises the day in two minutes. Free, no spam, unsubscribe anytime.