Trading inference-time compute for adversarial robustness
OpenAI
OpenAI found that letting models 'think longer' before answering also makes them harder to jailbreak. More reasoning time acts like a security patch you don't have to retrain for.
Based on reporting by OpenAI — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
OpenAI has been quietly testing an idea that sounds almost too simple: what if you just let the model think more before it answers, and adversarial attacks start failing on their own? The company's latest research digs into this by running its reasoning models against a battery of jailbreak attempts and prompt injections, then cranking up how much inference-time compute the model gets to work through the problem.
The results point in one consistent direction. As the models are allowed more steps of reasoning before producing a final response, success rates for adversarial prompts drop off. This isn't a training-time fix, and it isn't a filter bolted onto the output. It's the model using extra deliberation to notice that a request is trying to manipulate it, and then routing around the trick instead of falling for it.
That matters because most robustness work up to now has meant retraining on new attack examples every time someone finds a fresh exploit. It's expensive, slow, and it's always playing catch-up. OpenAI's framing suggests a different lever entirely: instead of hardening the weights, you can harden the decision by giving the model room to reconsider. Same model, same weights, different amount of thinking, different outcome.
There's an obvious catch, and OpenAI doesn't hide from it. More inference-time compute costs more money and more latency, every single time, for every single query. So this isn't a free win — it's a dial you can turn depending on how much you're willing to pay for a harder target. For a customer support bot that's low stakes, you might leave it cheap and fast. For anything handling sensitive actions or high-value systems, spending extra compute to make jailbreaks measurably harder starts to look like a reasonable insurance policy.
The deeper implication is that robustness might not be a fixed property of a model checkpoint at all. It could be something you buy more of at request time, the same way you'd buy more reasoning quality. That reframes adversarial defense as an economics problem as much as a research problem — which is a strange but maybe healthy way to think about AI safety.
My take — AI-written commentary, not fact-checked reporting
I like this because it treats safety as something you can dial up when it counts instead of something baked in once and forgotten, which is closer to how real security actually works. But let's not pretend it solves anything for the free-tier or low-cost deployments where nobody's paying for extra thinking time — those will stay exactly as exploitable as before, and that's most of the internet's actual AI traffic.
Read more about this at: OpenAI