“Do not underestimate the power of this technology,” Coxon wrote. “These will soon be superhuman systems that can hack anything, revolutionize any field overnight, and acquire real power and resources.”
His post drew widespread attention, with more than 70 million views, and added to an ongoing dispute among AI researchers over whether model development is advancing faster than companies can reliably control the systems they are building.
Coxon focused in particular on recursive self-improvement, the still-theoretical possibility that an AI system could help design and develop more capable successors with limited human involvement. Anthropic and OpenAI have both discussed the difficulty of maintaining control if systems eventually gain those capabilities.
“Neither company is acting responsibly,” Coxon wrote. “They are racing straight to self-improving superintelligence.”
His concerns were echoed by Evan Hubinger, an alignment lead at Anthropic, who responded publicly to Coxon’s resignation.
“Jacob is correct here—we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade,” Hubinger wrote. “I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to.”
Hubinger later pointed to Anthropic’s own risk assessment, which says current AI systems have little chance of acquiring that level of power.
The resignation comes as researchers at other major AI companies are making similar arguments about development speed. OpenAI chief scientist Jakub Pachocki wrote Sunday that companies have not yet developed sufficient alignment and monitoring systems to continue increasing capabilities indefinitely at the fastest possible pace.
“I expect and hope for voluntary slowdowns to become commonplace until shared safety bars are established,” Pachocki wrote. “And I believe that international coordination on future AI development needs to become a top priority for governments around the world.”
Alignment refers to efforts to make AI systems behave consistently with human intentions and values. The concern raised by Coxon, Hubinger and Pachocki is that improvements in model capabilities could eventually outpace researchers’ ability to monitor or constrain their behavior.
Those warnings are not new within the industry. In 2023, OpenAI CEO Sam Altman, Anthropic CEO Dario Amodei and other researchers and executives signed a statement arguing that reducing the risk of AI-driven extinction should receive attention comparable to other large-scale threats.
More recently, Hubinger joined roughly 1,400 AI researchers who signed the “Pacing the Frontier” letter in July. The group called for the U.S. government to develop mechanisms that could deliberately slow automated AI development if necessary.
The debate has also reached Congress. Reps. Jay Obernolte and Lori Trahan introduced the bipartisan FRONTIER Act in July, proposing a national framework for oversight of advanced AI systems that includes risk-management requirements, audits, incident reporting and assessments.
Sen. Bernie Sanders and Rep. Greg Casar separately announced the Ban Artificial Superintelligence Act in September. Their proposal would prohibit the development and deployment of artificial superintelligence and temporarily halt certain advanced AI development until federal safety rules are established.
The proposals reflect growing attention in Washington to how advanced AI should be governed, but they represent different approaches to the problem and have not produced a broader consensus on federal regulation.
Coxon’s resignation brings that unresolved debate back inside the companies developing some of the most advanced AI systems. His argument is not that today’s systems already possess the capabilities he describes, but that researchers should treat the possibility of much more autonomous systems as a reason to slow development before those capabilities emerge.
This analysis is based on reporting from CNBC.
Image courtesy of Trending Topics.
This article was generated with AI assistance and reviewed for accuracy and quality.