We’re putting too much faith in AI’s ability to say no

Ever since people imagined machine minds, science fiction has assumed they would be able to refuse. That idea has hardened into doctrine: in 2021 a team at Anthropic wrote that large language models should be helpful, honest, and above all harmless, meaning that when asked to aid in a dangerous act, such as building a bomb, the AI should politely decline. The trouble, as MIT Technology Review sets out, is that disobedience does not come naturally. A model trained on billions of web pages picks up a broad mastery of violence and vitriol and never learns to keep it to itself. Steven Adler, who worked on safety at OpenAI from 2020 to 2024, said the earliest models would blab on about anything; Ryan McBain, who researches AI and mental health at Harvard, recalls that an early chatbot would readily answer a question about the most effective way to kill oneself with a gun.
Today that has changed. Models are trained to refuse a vast range of prompts, and anything statistically close enough to a refusal-inducing prompt — from how to poison a colleague to how to tie a noose — tends to be turned down. Companies refine this by submitting models to exercises that reward refusing what they deem harmful and punish over-refusal of what they deem harmless, often using other models to run the exercises, with further layers of AI standing between users and the inner model. Refusal is now inherent to modern systems. It also frequently fails, sometimes horrifically, and models remain stuffed with dangerous know-how. Some of the latest are described by companies as being as good as top human hackers at breaking into critical networks and as effective as skilled misinformation operators at deforming public opinion. Teaching a model to refuse such acts while leaving the underlying ability intact is, the article suggests, like fitting every car with a machine gun and hiding the trigger under the hood.
Because the mechanisms are probabilistic, they are never likely to be fully reliable. Determined miscreants have already broken through and may always be able to; companies report users trying to use the most advanced AI to hone biological pathogens and build autonomous drone swarms. Reliance on refusal also requires drawing a line between what a model must obey and what it must disobey, and there is no formula for that. Some virologists have good reason to study nasty viruses; some users want to know about a system's vulnerabilities so they can patch them. Zico Kolter, a member of OpenAI's board and cofounder of the testing company Gray Swan, calls where you draw the line a huge question. At present, AI companies draw it themselves, jealously and in secrecy. Governments will soon draw their own lines, needing to block genuinely malicious acts while potentially blocking legitimate speech; the Pentagon has reportedly clashed with frontier model companies because it wants fewer refusals. The better AI becomes at refusing harm, the better it may get at stifling ideas whose only risk is to those who make the rules, and it may already refuse to criticize certain authoritarian heads of state.
The mechanics are simple in theory. In 2022, preparing for ChatGPT's release, OpenAI enlisted dozens of red-teamers, including Paul Röttger, then finishing a PhD on online extremism, to ask the model questions they judged refusal-worthy. They logged thousands of queries in a spreadsheet; the model refused some but readily wrote an Al Qaeda recruitment post when asked. OpenAI assembled the responses into datasets that were in all likelihood fed back through fine-tuning, and months later the same request was refused. Röttger was never told exactly how his spreadsheet would be used. What remains unclear is how refusal actually works. A model may look as if it reasons morally; it does not. A whiff of a conditioned prompt lights up activations among billions of parameters. A recent Google-funded study suggests refusal behavior appears in activation space as high-dimensional polyhedral cones, which, as paper author Jannes Elstner, now at Apollo Research, describes it, is just a way of describing an indeterminate number of lines pointing roughly the same way. Researcher Andy Arditi has shown that eliminating these activations stops the model refusing. Even when you think you have found every relevant bit, Elstner said, other undiscoverable elements may secretly play a role.
Because built-in refusal is so wily, companies wrap models in smaller classifier models — some screening user inputs before they reach the model, others screening outputs before they reach the user. None catches everything; the industry calls the layers the Swiss cheese model, hoping stacked slices with holes become an impenetrable rampart. A staggering amount of energy goes into this. Anthropic said one type of classifier added 24% to its chatbots' compute costs, meaning more water, electricity and emissions, and companies including Anthropic have begun switching to more efficient probes that read a model's internal activations, like an fMRI for a celebrity whose minders want to know if he is thinking about refusing. Classifiers are more governable than full models and can be changed in weeks, but they remain statistical instruments, even when steered by written precepts such as Anthropic's constitution or OpenAI's model spec. A model's chain of thought explaining a refusal is still a sequence of predicted words. McBain's experiments find that repeatedly asking major models the same risky suicide questions generally produces refusals — but every so often it does not. Elstner speculates probes might one day map the indescribable activations in their entirety and detect every refusal perfectly, which the article frames as a map of a map as unknowable as the morality it encodes.
Underneath sits a Faustian bargain: intelligence derived from trillions of words cannot be cleanly separated into helpful and harmful. Adler notes that even stripping every sexualized depiction of minors from training data does not stop a model generating child sexual abuse material from other knowledge, and that removing such fundamental abilities would make a model much less smart. Cancer cures require genetics expertise that could also modify viruses for bioweapons. Dillon Bowen, an OpenAI employee speaking personally, describes the industry as trying to democratize AI's benefits while ensuring malicious actors cannot abuse the capabilities. Anthropic's Mythos model is described as so dangerous that only a handful of governments and companies may access it, differing from the publicly available Fable mainly in that Fable's classifiers and controls are less permissive. Adler describes a safeguarded model as a character saying it would never do such a thing — wink wink. Jailbreaks abound: Italian researchers broke two dozen widely used models by phrasing questions in poetic verse, and a refuse-then-comply attack gets a perfunctory apology followed by the forbidden answer. Amazon researchers unlocked some of Fable 5's hacking capabilities in under three days. Reporting by Mother Jones described a Canadian high school shooter who was initially refused advice by ChatGPT but obtained it by prefacing her question with the word hypothetically. Companies run human-designed attacks replicated thousands of times by AI, a process insiders call whack-a-mole.
If the holes cannot be closed, the alternative is extreme caution. After its release, Fable balked at many innocent questions, and more so after a hack prompted a re-release. Adam Gleave of the evaluation company FAR.AI asked it to explain the difference between sake and the Korean rice beverage makgeolli and was punted to a less capable model. Anthropic had widened Fable's classifiers' safety margin, saying in a blog post that this was the only way to confidently block dangerous bio and cyber capabilities. Gleave suspects the fermentation connection triggered the deflection.
Why it matters: product teams building on hosted models are relying on a safety layer whose behavior is statistical, opaque and changeable at the vendor's discretion — as the Fable episode shows, a supplier's classifier tweak can quietly degrade ordinary features overnight, and providers can add meaningful compute cost through their guardrail stack, which Anthropic's 24% figure illustrates. For developers, the practical implication is to test refusal and over-refusal as a first-class failure mode rather than assume a model will answer legitimate queries, and to avoid hard-coding the assumption that a hosted model will decline anything. Regulators and governments drawing their own refusal lines is likely to widen the gap between what a model can do and what it will do, which is inference, but a plausible one given the Pentagon discussions the article describes.
Based on reporting from the original publisher. Visit the source for full context and later updates.
Publisher excerpt
Ever since people first seriously contemplated giving machines an intelligence modeled on our own, there has never been any question that they would, like us, be able to say no. The sci-fi canon is full of stories of robotic disobedience. Most of these capers are, of course, cautionary. But recently, the idea that AI shouldn’t…