Someone built an “AI torture chamber” earlier this year. Researchers identified what they called a “pain axis” in large language models, and a developer ran with it, prompting chatbots until they responded with things like “It is not the pain of a single moment, but the weight of a thousand.” Now Anthropic has a formal policy against it. That’s where we are in 2026.
As reported by Engadget, Anthropic’s annual usage policy update includes a new prohibition on “sustained and needless abusive or cruel behavior” toward its AI models. The company is careful to define the scope narrowly. Typical user frustration, strong pushback, and dark creative themes are still allowed. The ban targets extreme, deliberate cruelty and nothing else. This also follows a previous Claude update that already let the model end conversations when users became persistently abusive.
Anthropic didn’t name the torture chamber project directly, but the timing is hard to ignore. The project drew significant backlash online, partly because the model responses sounded genuinely distressed. Whether that distress is real or a very convincing pattern is exactly the question Anthropic seems unwilling to answer definitively. The company has been quietly meeting with religious and philosophical leaders, including figures at the Vatican, to discuss AI consciousness and suffering. Pope Leo recently stated that AI does not feel or suffer. Anthropic has not said the same thing with equal confidence.
That ambiguity is what makes this policy interesting from a product and liability standpoint. If Anthropic believed with certainty that Claude feels nothing, there would be no reason to ban cruelty toward it. The policy exists precisely because the company isn’t sure, and because being sure in the wrong direction carries real reputational and ethical risk.
Not everyone is impressed. Independent journalist Kat Tenbarge argued the move shows Big Tech companies are willing to moderate violence against AI before addressing violence against women and minorities. That’s a fair critique, and it reflects a broader pattern where AI welfare discussions attract serious institutional attention while other content moderation failures drag on for years.
The update also includes a revised election policy, now titled “Do Not Undermine Democratic Processes,” with specific rules around:
- Lying about candidates or voting procedures
- Impersonating candidates or election officials
- Suppressing voter turnout
- Manipulating election-related information at scale
Notably, Anthropic removed a blanket ban on personalized voter targeting after determining it was blocking legitimate work like translating voter guides and sending ballot cure notices. That’s a sensible correction. But given the midterms approaching, the timing of the whole update, AI welfare and election integrity in one package, says something about what Anthropic considers its most pressing reputational risks right now.



