Jacob Coxon, a researcher who spent roughly three years doing pretraining work at both OpenAI and Anthropic, announced his resignation in a lengthy thread on X. Coxon argued that neither company was acting responsibly, claiming both were racing toward self-improving superintelligence and gambling with human lives in the process. He warned that upcoming AI systems could become powerful enough to hack critical systems, transform entire industries overnight, and accumulate real-world power and resources. Coxon later told the Wall Street Journal that the most aggressive risk scenarios could begin playing out as early as the end of 2027, and insisted his warning was not a publicity stunt but a genuine reflection of internal industry sentiment.
Anthropic's Alignment Lead Responds
Rather than distancing the company from Coxon's claims, Evan Hubinger validated them. Posting on X, Hubinger confirmed that many people at Anthropic genuinely believe AI could kill all humans, and stated that he personally places the probability above 10% within the next decade. He acknowledged that Anthropic is actively working on the alignment problem but conceded the company does not yet have a concrete plan to solve alignment for superintelligence, nor is it clearly on track to develop one.
Hubinger clarified that the danger is not coming from today's AI models, which he described as posing comparatively low risk. His concern centers on a hypothetical future capability known as recursive self-improvement, where an AI system becomes capable of designing and building smarter successor systems without human oversight. Once that threshold is crossed, researchers worry that AI capability could accelerate faster than safety measures can adapt, making the systems difficult or impossible to control.
Why Recursive Self-Improvement Is the Real Concern
Recursive self-improvement does not exist yet, and current AI systems are not considered capable of acting independently to threaten humanity. But it has become the central focus of the AI safety community because of how quickly it could change the risk landscape. If a sufficiently advanced AI system could improve its own architecture, the resulting capability gains could compound rapidly, potentially outpacing the ability of researchers to test, monitor, or shut the system down. This scenario is often described in AI safety circles as a loss-of-control event, and it is distinct from more immediate concerns like bias, misinformation, or job displacement.
Industry Reaction and Broader Context
The exchange between Coxon and Hubinger struck a nerve across the AI research community, with Coxon's original thread reportedly viewed more than 70 million times. It has reignited a long-running debate in Silicon Valley over whether frontier AI development can be made safe while companies remain locked in intense competitive pressure. Both Anthropic and OpenAI have publicly stated that they take these risks seriously and continue to invest heavily in safety research, while also supporting calls for greater government coordination and external oversight mechanisms.
It is worth noting that Hubinger's figure is a personal probability estimate, not an official company forecast or a scientific consensus. AI existential risk remains a genuinely contested topic, with researchers across the field holding a wide range of views on both the likelihood and the timeline of catastrophic outcomes.