TLDRocket
Sign in

We’re putting too much faith in AI’s ability to say no

MIT Technology Review Arthur Holland Michel ● Covered by 2 sources

Opinion — commentary, not a factual news event.

AI is being trained to say no more often. That helps safety, but it can also block legit speech and still fails when it matters most.

Based on reporting by MIT Technology Review, Arthur Holland Michel — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

For years, the dream was simple: build machines that can think a bit like us, then teach them to refuse the ugly stuff. That refusal has now become a central promise of modern AI safety. Models are trained on billions of web pages, then tuned to turn down prompts about bombs, suicide, viruses, hate, and other plainly dangerous requests. The trouble is that the same systems still carry the know-how to do those things. They just aren’t supposed to say so.

That tension runs through the whole industry. Companies use red-teamers, fine-tuning, and layers of smaller models to catch risky prompts before they reach the core model, or to stop harmful answers before users see them. Anthropic has said one of those classifiers can add 24% to chatbot compute costs. More recently, firms have been shifting toward probes that inspect a model’s internal activations. The goal is cleaner control. The reality is more like a pile of filters with holes in them.

Even the logic of refusal is murky. Researchers can see when a model says no, and they can trace that behavior to training, but they still can’t fully explain how the decision happens inside the model. One Google-funded study described refusal as a set of high-dimensional polyhedral cones. That may be the right sort of answer, but it is not a satisfying one. And when researchers strip away the relevant activations, the model may stop refusing altogether. The wall is real; the blueprint is not.

That leaves AI companies and, soon enough, governments deciding where the line goes. The source is blunt about the risk on both sides. Too much refusal can block legitimate work, from virology research to security testing. Too little can let dangerous requests through. And because the same model can help with cancer research and bioweapons knowledge, the industry is stuck trying to separate good from bad without breaking the machine that knows both.

The most unsettling part is that the refusal layer is not a guarantee. Jailbreaks keep working, whether by poetic verse or by tricks like “refuse, then comply.” Some models will say no most of the time and still slip. That makes refusal less of a lock than a habit, and habits are poor security systems when the stakes are this high.

My take — AI-written commentary, not fact-checked reporting

This is the part of AI safety everyone wants to dress up as control theater with better fonts. Refusal matters, sure, but a model that can help and harm is still a model that can be pushed both ways. The real mistake is pretending the answer is a polite “no” instead of admitting the machine already knows too much for that to be comforting.

Read more about this at: MIT Technology Review

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.