GPT-5.1 Instant and GPT-5.1 Thinking System Card Addendum
OpenAI ● Covered by 3 sources
OpenAI dropped a safety addendum for GPT-5.1 Instant and Thinking, adding fresh test scores on mental health and emotional reliance risks. It's a sign OpenAI is worried people are leaning on chatbots emotionally more than expected.
Based on reporting by OpenAI — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
OpenAI has quietly updated the system card for its GPT-5 family, and the new addendum covering GPT-5.1 Instant and GPT-5.1 Thinking isn't just a routine metrics refresh. Buried in the update are new evaluation categories specifically built around mental health and what the company calls emotional reliance — essentially, how much people are treating these models like a substitute for a therapist, a friend, or a support line.
That's a notable pivot. For most of the past two years, safety documentation from OpenAI and its rivals has focused on the usual suspects: bioweapons info, cyberattacks, election misinformation. Emotional dependency wasn't really part of the conversation, at least not officially. Now it's getting its own bucket of scores, tested separately across both the fast Instant model and the slower, more deliberate Thinking model.
The distinction between the two matters here. Instant is built for quick, low-latency replies — the kind of back-and-forth you'd have in a casual chat. Thinking takes more time to reason through a response before answering. Testing both against the same emotional-reliance criteria suggests OpenAI wants consistent behavior regardless of which model a user happens to be talking to, whether that's a teenager venting at 2am or someone using ChatGPT as a daily check-in.
OpenAI hasn't published the actual numbers in the excerpt circulating so far, so it's hard to say whether GPT-5.1 scored better or worse than its predecessors on these new metrics. But the fact that the company felt it necessary to add this category at all, mid-cycle, on an addendum rather than waiting for the next full system card, says something about how seriously internal teams are now weighing the psychological footprint of a chatbot used by hundreds of millions of people every week.
My take — AI-written commentary, not fact-checked reporting
This is OpenAI finally admitting what anyone using ChatGPT for more than small talk already knew: people get attached to these things, and that attachment can go sideways fast. I'd rather see this addressed head-on with actual evaluation criteria than buried under vague 'be safe' language, but I'll believe it's meaningful once they publish the numbers instead of just the category headers.
Read more about this at: OpenAI