SafetyKit scales risk agents with OpenAI’s most capable models
OpenAI
SafetyKit is now running its trust-and-safety agents on OpenAI's GPT-5 to catch policy violations faster and more accurately. Old-school moderation tools built on rigid rules just can't keep up with how fast platforms and bad actors evolve.
Based on reporting by OpenAI — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
Content moderation has always been a game of whack-a-mole. Platforms write rules, bad actors find loopholes, platforms patch rules, repeat forever. SafetyKit, a company that builds automated risk and compliance agents for online platforms, is betting that OpenAI's newest models change the math on that cycle.
The pitch is straightforward: swap brittle keyword filters and hand-coded decision trees for GPT-5-powered agents that can actually reason about context. A post that mentions a drug name isn't automatically a violation — it depends on whether it's a harm-reduction resource, a marketplace listing, or a joke. Legacy systems tend to miss that nuance or overcorrect into false positives, and both failure modes cost platforms money, either through fines and enforcement actions or through banning legitimate users.
What's notable here isn't just that SafetyKit adopted a bigger model. It's that they're using it to scale agents that make judgment calls at a speed no human moderation team could match, while claiming accuracy gains over the rule-based systems that have dominated trust and safety for years. OpenAI, for its part, is happy to showcase this as proof that GPT-5 isn't just a chatbot upgrade — it's infrastructure for decisions that used to require a person reading a screen.
The compliance angle matters too. Platforms in regulated spaces — marketplaces, fintech, health-adjacent apps — face growing pressure from regulators to prove they're actively policing their own content, not just reacting after the fact. An AI agent that can flag, explain, and act on violations in near real time is a much easier thing to point to during an audit than a spreadsheet of manual review logs.
Whether this holds up outside a case study is the real question. Model-based moderation still hallucinates, still gets manipulated by adversarial prompts, and still needs human oversight when the stakes are high. SafetyKit's bet is that GPT-5 is good enough, often enough, to make that oversight the exception rather than the rule.
My take — AI-written commentary, not fact-checked reporting
I'll believe the accuracy claims when someone publishes the false-positive and false-negative rates instead of a vendor case study, because every moderation vendor says their new model is better right before the next embarrassing screenshot goes viral. That said, this is the correct direction — dumb keyword filters have been failing users and platforms for a decade, and outsourcing judgment calls to a smarter model beats pretending rigid rules can cover every edge case. Just don't remove the humans yet.”
Read more about this at: OpenAI