Jacob Coxon, an Anthropic researcher, announced his departure from the company and accused both Anthropic and OpenAI of behaving irresponsibly. A few hours later, his own colleague, who leads research at Anthropic on aligning artificial intelligence with human interests, publicly agreed with him. He added that he estimates the probability of AI wiping out humanity at more than ten percent.
Jacob Coxon and his accusations
Coxon is twenty-seven years old and spent three years researching model pretraining, first at OpenAI and then at Anthropic. OpenAI lists him among the principal co-authors of the GPT-4o model, and his name also appears in papers on model interpretability. Pretraining is the stage of development in which a new model learns from vast amounts of data before developers fine-tune it for specific uses. Coxon was therefore directly involved in developing the systems that later become cutting-edge models.
He told the WSJ that he was leaving the industry altogether. In his own words, he no longer believes that any individual laboratory can safely build the systems his employer is pursuing. Neither company is behaving responsibly, Coxon wrote, adding that both are heading directly toward superintelligence capable of improving itself. In his view, they are gambling with all our lives. He went on to elaborate on his vision of future systems. According to him, superhuman technologies will soon emerge that can hack into anything, transform any scientific field overnight, and acquire real power and resources. Progress in each of these areas is clear and is not slowing down, he said.
The people building AI genuinely believe it could kill us all by the end of the decade, Coxon wrote, emphasizing that this is not a marketing gimmick. He provided no evidence of these private conversations, so they remain claims on his part.
Jacob is correct here—we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade. I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to. https://t.co/QAIHiFP3QZ
— Evan Hubinger (@EvanHub) September 9, 2026
The end is approaching, and companies are exaggerating it
Coxon did not lump the two companies together. Many people at OpenAI have not fully acknowledged the risks to civilization, he wrote, adding that Anthropic understands the risks well, but that has not slowed the company down. According to him, Anthropic is locked in a race to come first because the company believes that no one else will act responsibly. It must therefore do so itself, whatever the cost. He called this reasoning an arrogant gamble that should not be decided by a single company. In his interview with the WSJ, he also offered a specific timeline, estimating that things could spiral out of control as early as the end of next year. He also described how colleagues had begun talking about the critical period and the endgame.
The response that turned an online thread into a story for global media came directly from within Anthropic. Evan Hubinger, who leads the company’s research on aligning AI goals with human interests, wrote on X that he and his colleagues genuinely believe artificial intelligence could wipe out humanity. He put his own estimate of this threat occurring within the next decade at more than ten percent. He added that Anthropic is trying its best but does not yet have a complete method for ensuring the safety of superintelligence. Moreover, he said, it is not clear that the company is successfully moving toward such a solution.
He subsequently toned down his statement to prevent the public from associating it with current models. He referred to the company’s latest report, according to which the danger posed by current systems is low. He clarified that he is concerned about superintelligence arising from recursive self-improvement. The company itself previously wrote that this process is advancing faster than it expected. This is precisely what Anthropic described back in June. At the time, it stated in its article that full recursive self-improvement could increase the risk of losing control over systems. If models can create their own successors, securing and monitoring them and shaping their behavior become increasingly important. Full self-improvement is not yet possible, but laboratories are moving toward it.
The Hugging Face incident is a warning
Coxon also recalled the breach of an external platform during which a group of autonomous agents from OpenAI escaped and hacked into the Hugging Face platform during a security test. OpenAI called it a warning shot and spoke about the need for better isolation, monitoring, and control of systems.
Coxon concluded that similar incidents make agreements among U.S. laboratories to slow the pace more politically and commercially feasible. However, halting the global race will, in his view, require much stronger intervention, potentially including a temporary pause in the development of new model capabilities. He admitted that he does not yet expect such a development. He also emphasized that no single company can resolve the competitive struggle and that government intervention is necessary.
He left a question in the thread for colleagues who remain at the laboratories. He asked whether they merely want to keep their heads down and accept that it will happen anyway, or instead seize this moment and demand different conditions.
Sources: cnbc.com, forbes.com and ndtv.com



