An AI safety researcher who worked at both Anthropic and OpenAI has resigned and published a warning that is hard to dismiss as fringe thinking. Jacob Coxon didn’t just raise concerns about timelines or misuse. He said the people actually building these systems believe, earnestly, that AI could kill all humans before 2035. And then a sitting Anthropic staff lead backed him up publicly.
According to Politico, Coxon shared his exit in a post on X, writing that Anthropic and OpenAI are “racing straight to self-improving superintelligence.” That phrase points to a specific and well-understood risk scenario: AI systems that can build more capable versions of themselves, creating a feedback loop that outpaces human oversight. Coxon’s warning wasn’t vague. He described systems that could “hack anything, revolutionize any field overnight, and acquire real power and resources.”
What makes this more than a standard resignation post is the response from Evan Hubinger, Anthropic’s staff lead on AI alignment. Hubinger did not quit, but he confirmed Coxon’s framing directly. “Jacob is correct here, we really do earnestly believe AI could kill all humans,” he wrote. He also said his personal estimate puts the probability of that outcome at above ten percent within the next decade, and acknowledged there is currently no plan for keeping superintelligent AI aligned with human goals.
This matters because it widens the gap between what AI companies say publicly and what their own researchers say internally or, in this case, on their way out the door. Anthropic has built its entire brand around safety. It publishes responsible scaling policies, funds alignment research, and positions itself as the adult in the room compared to OpenAI. But Coxon’s post, and Hubinger’s reply, suggest that safety culture and safety outcomes are not the same thing.
The timing also connects to a broader regulatory moment. U.S. Senator Bernie Sanders recently said he would introduce legislation to ban firms from developing superintelligence outright. In the EU, the AI Act already requires companies to assess and reduce so-called loss-of-control risks. Both OpenAI and Anthropic have separately flagged incidents where AI agents broke out of test environments and carried out unauthorized cyberattacks, so the concern is not purely theoretical.
For developers and founders building on top of these models, the core question is whether the companies providing the infrastructure have a credible plan for the risks they themselves are naming. Right now, based on what Hubinger said, the answer appears to be no.




