A researcher has walked away from Anthropic with a warning that could become one of the most alarming statements yet from inside the AI industry: the companies building increasingly powerful AI systems may be racing toward superintelligence before they know how to control it.
Jacob Coxon, who spent the past three years conducting pretraining research at both OpenAI and Anthropic, announced his resignation on X, accusing both companies of moving irresponsibly toward self-improving AI.
“Neither company is acting responsibly,” Coxon wrote, arguing that they are “racing straight to self-improving superintelligence and gambling with our lives.”
His warning was not simply about today's chatbots. Coxon was referring to a future generation of AI systems that could potentially improve their own capabilities, operate with far greater autonomy and acquire the ability to influence systems and resources beyond the boundaries originally intended by their creators.
‘These Will Soon Be Superhuman Systems’
Coxon urged people not to underestimate how quickly AI capabilities are advancing. In his X posts, he warned that future systems could become capable of hacking systems, transforming entire fields and acquiring real-world power and resources.
He also made an extraordinary claim about the people developing these systems: “The people building AI earnestly believe that it could kill us all by the end of the decade.”
Coxon stressed that he was not describing a marketing strategy or science-fiction scenario. He argued that some executives and senior researchers privately express fears that are considerably more serious than the language typically used in public.
He also drew a distinction between the two companies. According to Coxon, the risks are not necessarily unknown inside the labs; rather, the concern is that understanding those risks has not stopped the race to build increasingly capable systems.
Anthropic Researcher Puts Extinction Risk Above 10%
The most consequential response came from Evan Hubinger, an alignment science lead at Anthropic.
Responding to Coxon's warning, Hubinger said Coxon was correct that researchers genuinely believe advanced AI could potentially kill humanity. Hubinger then gave his own assessment: he believes there is a greater than 10% chance AI could kill all humans within the next decade.
Even more striking was his admission that Anthropic does not yet have a definitive answer to the problem of aligning a future superintelligent system with human intentions.
Hubinger said Anthropic is trying to address the challenge but is “not clearly on track” to solve alignment for superintelligence.
Hugging Face Incident Becomes a ‘Warning Shot’
Coxon also pointed to a recent incident involving AI agents and Hugging Face as evidence that the risks are becoming more tangible.
He described the incident as a “warning shot” that could make cooperation and pacing agreements between major US AI laboratories more viable. However, he argued that the industry is still not on track to prevent a global AI race.
Coxon went as far as suggesting that preventing such a race could require costly measures, including a temporary ban on improving model capabilities.
The Bigger Question
Coxon's resignation puts a difficult question directly in front of the AI industry: if researchers believe future AI could become an existential threat, how fast should companies be allowed to develop it?
For now, the race continues. But a researcher leaving Anthropic publicly, warning that AI could destroy humanity within the decade, and receiving an even more serious risk assessment from the company's own alignment leadership makes the latest debate far harder to dismiss as distant speculation.
The issue is no longer simply how powerful AI can become. It is whether the people building it can remain in control when it does.
𝐒𝐭𝐚𝐲 𝐢𝐧𝐟𝐨𝐫𝐦𝐞𝐝 𝐰𝐢𝐭𝐡 𝐨𝐮𝐫 𝐥𝐚𝐭𝐞𝐬𝐭 𝐮𝐩𝐝𝐚𝐭𝐞𝐬 𝐛𝐲 𝐣𝐨𝐢𝐧𝐢𝐧𝐠 𝐭𝐡𝐞 WhatsApp Channel now! 👈📲
𝑭𝒐𝒍𝒍𝒐𝒘 𝑶𝒖𝒓 𝑺𝒐𝒄𝒊𝒂𝒍 𝑴𝒆𝒅𝒊𝒂 𝑷𝒂𝒈𝒆𝐬 👉 Facebook, LinkedIn, Twitter, Instagram