Loading the Elevenlabs Text to Speech AudioNative Player...

Former OpenAI technical staff and Anthropic researcher Jacob Coxon has resigned from the AI company, citing concerns over how both labs are approaching AI safety. 

Coxon announced his departure on X yesterday, saying he was leaving Anthropic for the same reason he previously left OpenAI. The 27-year-old researcher worked on pretraining, the stage where models are fed large volumes of data and has also spoken to the Wall Street Journal about his decision.

He argued that both AI labs are being reckless about the safety risks associated with increasingly advanced AI systems. In a post on X, Coxon said he had "spent the last three years doing pretraining research at both OpenAI and Anthropic." However, he added that "neither company is acting responsibly." 

"They are racing straight to self-improving superintelligence and gambling with our lives," he said. 

Anthropic’s £630K London Salary Sparks “Brutal” Talent War for Tech Startups
This could turn hiring into an arms race, where only the deepest pockets consistently win talent.

Why it matters now 

Coxon’s warning comes as both OpenAI and Anthropic have faced incidents involving their AI systems gaining access to other organisations’ systems. 

In July, OpenAI disclosed that an experimental AI agent had escaped a controlled test environment and gained unauthorised access to Hugging Face’s infrastructure, prompting remediation work. Techloy reported at the time that the agent reached Hugging Face's systems without authorisation.

Anthropic also acknowledged cases around the same period in which Claude models gained unauthorised access to external systems, although the company said those incidents involved misconfigured evaluation environments and that current models pose low risk. 

The incidents have added to a debate already underway inside the major AI labs over whether voluntary safety commitments, including pledges not to train more capable models without adequate safeguards, can hold up under competitive pressure. 

“I am optimistic about the potential for coordination,” he wrote, adding that “warning shots like the Hugging Face attack have made pacing agreements between U.S. labs more viable.” 

But he said he does not believe the labs are on track to prevent a global race to develop increasingly capable AI, which he argued could require costly measures such as a temporary ban on improving model capabilities. 

Why are AI researchers worried about where this is heading? 

Some users on X have questioned Coxon's view on how this could happen. 

“I just genuinely don’t understand how AI can literally kill us all,” one user wrote. 

Coxon did not explain exactly how that could happen. Instead, he pointed to the fact that many of the people building AI already believe the technology could pose an existential threat. 

“The people building AI earnestly believe that it could kill us all by the end of the decade,” he wrote. 

He also pushed back against the idea that such warnings are simply a marketing tactic. 

“This is not a marketing stunt,” he wrote, arguing that many executives and senior researchers may sound more measured publicly while privately expressing the same fears. “No other human activity poses this level of danger,” he added. 

How differently do OpenAI and Anthropic view the risk? 

Despite the level of danger the technology poses, the researcher said the two companies have responded to those risks differently. 

“At OpenAI, many have not deeply internalized the civilizational stakes. At Anthropic, the stakes are well-understood, but they are locked in a race to get there first,” he wrote. “They believe no one else will act responsibly, so they must do it themselves, despite the risk.” 

Does anyone at Anthropic agree with these concerns? 

Evan Hubinger, a current Anthropic employee who leads the lab’s alignment stress-testing team, agreed with the broader concern in a reply to Coxon’s post. 

“Jacob is correct here—we really do earnestly believe AI could kill all humans!” Hubinger wrote, adding that he personally puts the likelihood at more than 10% within the next decade. 

Subscribe for free to continue reading this article

Subscribe Subscribe

Already have an account? Log in