TLDRocket
Sign in

Introducing Lockdown Mode and Elevated Risk labels in ChatGPT

OpenAI

OpenAI just added a Lockdown Mode and Elevated Risk labels to ChatGPT. It's meant to stop sneaky prompt attacks from stealing your data.

Based on reporting by OpenAI — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

OpenAI is rolling out two new defensive features for ChatGPT aimed squarely at organizations worried about prompt injection and data leaks. The first, Lockdown Mode, gives admins a way to clamp down on what the assistant can do when it's handling sensitive tasks or connected to external tools. The second, Elevated Risk labels, flags interactions or content that carry a higher chance of manipulation, so teams can spot trouble before it spreads.

Prompt injection has quietly become one of the thorniest problems in deploying AI agents at scale. Attackers don't need to hack a model directly. They just need to hide instructions inside a webpage, a document, or an email that the AI later reads and, without meaning to, obeys. That's how data exfiltration happens: not through a broken firewall, but through an assistant convinced it's following orders from its actual user.

What's notable here is that OpenAI isn't just patching a model weakness with more training data. Lockdown Mode is a control layer, something enterprises can toggle on for high-stakes workflows, effectively telling ChatGPT to distrust more of what it encounters and refuse riskier actions by default. Elevated Risk labels work more like a warning system, giving security teams visibility into which conversations or content sources look suspicious rather than leaving them to find out after the fact.

This fits a pattern that's been building all year. As companies wire ChatGPT and similar tools into internal systems, customer data, and third-party APIs, the attack surface stops being about jailbreaking the model and starts being about what the model can reach. OpenAI's move suggests it's treating that reach, not just the model's raw capability, as the thing that needs guardrails now.

My take — AI-written commentary, not fact-checked reporting

This is OpenAI finally admitting that capability without containment is a liability, not a feature. I'd rather see labels and lockdowns built into the product than another blog post promising the model is 'aligned' and hoping enterprises don't ask what happens when it reads a poisoned PDF. Still, calling something Elevated Risk after the fact doesn't fix an architecture that trusts untrusted input by default — that's the real fix nobody's shipped yet.

Read more about this at: OpenAI

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.