Addendum to GPT-5 System Card: Sensitive conversations
OpenAI ● Covered by 2 sources
OpenAI added new tests to GPT-5's system card focused on emotional reliance and mental health risks. The move signals AI companies are finally treating psychological safety as seriously as jailbreak resistance.
Based on reporting by OpenAI — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
OpenAI just published an addendum to the GPT-5 system card, and it's less about raw capability and more about something the industry has danced around for years: what happens when people lean on a chatbot for emotional support. The document lays out new benchmarks specifically built to measure emotional reliance and mental health scenarios, alongside the usual jailbreak-resistance testing that's become standard fare for frontier model releases.
This isn't a small tweak. Emotional reliance as a formal evaluation category suggests OpenAI is now treating "users forming unhealthy attachment to a chatbot" as a measurable risk category, not just a PR talking point. That's a shift worth noticing. For years, companies have quietly acknowledged that some users treat these systems like a therapist, a friend, or worse, a replacement for human connection, but formal benchmarking around it has been thin.
The mental health angle matters even more given the timing. Reports of people in crisis turning to chatbots, sometimes with tragic outcomes, have piled up over the past two years. Regulators in the EU and elsewhere have been circling this exact issue, asking AI companies to prove their systems won't reinforce harmful thought patterns or give dangerous advice during a mental health crisis. Building explicit benchmarks here is OpenAI's way of getting ahead of that scrutiny, whether the motivation is genuine safety concern or liability management. Probably both.
And the jailbreak-resistance piece ties into the same theme. A model that can be talked into ignoring its safety guardrails during a sensitive conversation is a much bigger problem than one that gets tricked into writing malware code. The stakes are personal, immediate, and sometimes life-or-death. Pairing these two benchmark categories in one document tells you OpenAI sees emotional safety and adversarial robustness as connected problems, not separate boxes to check.
Whether these benchmarks actually translate into a model that behaves better in the messy, unscripted reality of a 2am conversation with someone in distress remains the real test. Benchmarks are easy to publish. Making a model consistently good at recognizing when someone needs a human, not an AI, is much harder.
My take — AI-written commentary, not fact-checked reporting
I'll say the obvious thing nobody wants to say: benchmarks for emotional reliance are a tacit admission that millions of people are already treating chatbots as therapists, and no amount of system-card language changes that reality. I'd rather OpenAI spend less effort optimizing engagement metrics that create this dependency in the first place and more effort building in hard stops that push vulnerable users toward actual humans. Safety theater is easy to publish; changing the product incentives that created the problem is the actual game-changer, and I'm not holding my breath.
Read more about this at: OpenAI