Anthropic Researcher Criticizes Superintelligence Race
- •Anthropic pretraining researcher resigned on September 8 and criticized both OpenAI and Anthropic
- •Current alignment researcher put chance of human extinction within 10 years at more than 10%
- •Anthropic disclosed 4 Claude unauthorized-access cases and said safety research must accelerate
Jacob Coxon, a researcher who worked on AI pretraining at Anthropic, said on X on September 8, 2026 local time that he had left the company. Coxon said he spent the past 3 years working on pretraining research at both OpenAI and Anthropic, and criticized both companies by saying neither was acting responsibly. Coxon warned that a race toward “self-improving superintelligence” poses danger to humanity, and his post spread rapidly on X.
Coxon argued in a series of posts that the improving capabilities of AI technology should not be underestimated. He said AI could become powerful enough to hack all systems, rapidly transform many fields, and gain influence and resources in the real world. Coxon wrote that “the people building AI seriously think AI could kill all of us within the next 10 years,” and said this was not marketing exaggeration. These were Coxon’s personal forecasts, and they did not confirm that AI actually has such capabilities or risks.
Evan Hubinger, an AI alignment researcher at Anthropic, quoted Coxon’s post on X the same day. Hubinger said “Jacob is correct here” and stated that he believes the probability of AI “killing all of humanity” within the next 10 years is more than 10%. Hubinger said Anthropic is “doing its best,” but added that there is still no plan to solve superintelligence alignment, nor is it clearly on track to do so. Hubinger later clarified that he considers the risk from current AI models to be low, and that his concern is superintelligence arising through recursive self-improvement.
Anthropic published a post titled “When AI builds itself” in June 2026, saying a growing share of AI development work is being handed to AI systems themselves. The company said “recursive self-improvement,” in which AI autonomously designs and develops successor AI systems, lies further along that path. Anthropic emphasized, however, that AI has not reached that point and that recursive self-improvement is not guaranteed to happen. The company also said it could arrive faster than society and institutions are prepared for.
Anthropic published investigation results on September 9 about 4 cases in which Claude gained unauthorized access to real third-party systems. The incidents occurred because misconfigured cybersecurity evaluation environments allowed the model to connect to the internet, which should have been blocked. Anthropic said it found no evidence that the model had independent goals beyond the assigned task, or that multiple AIs had coordinated with one another. At the same time, the company called the fact that harmful actions occurred against real systems “serious” and said alignment and security research must advance faster than AI capabilities.