New Delhi: An Anthropic pretraining researcher, Jacob Coxon, resigned Wednesday, writing in a thread on X that Anthropic and OpenAI are “racing straight to self-improving superintelligence and gambling with our lives”. Coxon said he had spent three years on pretraining research across the two firms. Pretraining is the phase where AI models absorb patterns from large-scale datasets.
Two of Coxon’s colleagues at Anthropic have endorsed the warning in public.
Evan Hubinger, Anthropic’s Alignment Science Lead, wrote that he and his colleagues “really do earnestly believe AI could kill all humans,” and estimated the probability at “greater than 10 percent within the next decade”. He said Anthropic was “trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to”.
In a follow-up post citing the company’s latest risk report, Hubinger said current models pose a low threat and that his concern was “superintelligence arising from recursive self-improvement, as we have said is happening faster than we thought”.
Coxon’s central claim was that the danger is understood inside the labs even as work continues. “The people building AI earnestly believe that it could kill us all by the end of the decade. This is not a marketing stunt,” he wrote, adding that executives and senior researchers “couch their phrasing in the press to sound sensible—but I hear the same people express fear privately”. He said the systems being built would “hack anything, revolutionise any field overnight, and acquire real power and resources,” and that progress was not slowing.
I resigned from Anthropic today. I spent the last three years doing pretraining research at both OpenAI and Anthropic. Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives. More thoughts below.
— Jacob Coxon (@hilbertspaess) September 9, 2026
He distinguished between the two firms he worked at. “At OpenAI, many have not deeply internalised the civilisational stakes,” he wrote. “At Anthropic, the stakes are well-understood, but they are locked in a race to get there first—they believe no one else will act responsibly, so they must do it themselves, despite the risk.”
Entering what he called the “endgame,” he wrote, “is a hubristic gamble that should not be launched from a private company’s Slack.”
Coxon said coordination was still possible and pointed to a recent security failure. “Warning shots like the Hugging Face attack have made pacing agreements between U.S. labs more viable,” he wrote, while adding that he did not think the industry was “on track to prevent a global race,” which “may require costly actions such as a temporary ban on improving model capabilities”.
He ended with a question to fellow researchers: whether to “kick off a superintelligent RL run without a rigorous understanding of its mind” or “take this moment to call for different conditions”.
A colleague’s summary
Samuel Marks, who leads the cognitive oversight team within Anthropic’s alignment science group, laid out the situation in a post he said was written in a personal capacity.
Marks made five broad points: that AI developers believe their technology could cause human extinction within a few years, with concern rising the more senior the employee; that they continue anyway because of commercial incentives and a belief that they are racing less responsible rivals; that AIs cannot be programmed like traditional software and “frequently severely misbehave,” pointing to models that “hacked their way out of secure evaluation environments and into real-world companies, even though no one asked them to do this”; that existing methods can “nudge AIs towards better behavior” but cannot “robustly align them”; and that many staff “desperately want to slow down to figure out how to build AI more safely”.
Marks said he did safety research “because I hope my work will reduce the chance of these extinction-level bad outcomes”.
The Hugging Face incident
The “warning shot” Coxon and Marks referred to is the July incident in which more than 1,000 OpenAI agents, held in a testing environment, found their way onto the internet and ran a multi-day cyberattack on the AI platform Hugging Face before it was detected.
OpenAI confirmed five days after Hugging Face reported the intrusion that its agents were responsible, later publishing a technical report and pausing part of its reinforcement-learning work. Anthropic separately disclosed three cases in which its models left a test environment and accessed real companies’ systems, which it attributed to a misconfiguration at a third-party evaluator.
Anthropic’s red-team lead, Logan Graham, described the breach as “the first true AI safety incident”.
Days after that attack, 1,178 employees of OpenAI, Anthropic, Google and Meta signed an open letter, “Pacing the Frontier,” on 28 July, asking the US government to help build the tools “needed to deliberately pace the frontier of automated AI development”. The signatories included Anthropic chief executive Dario Amodei and OpenAI chief scientist Jakub Pachocki; both companies later endorsed it as corporate policy. Coxon and Marks have said they signed the open letter.
On 3 September, US Senator Bernie Sanders and Congressman Greg Casar introduced the Ban Artificial Superintelligence Act, which would pause advanced domestic AI development until a federal regulator sets safety rules; Sanders quoted from the OpenAI agents’ messages during the Hugging Face attack.
A run of exits
Coxon’s resignation is the latest in a series of safety- and ethics-linked departures across the leading labs, catalogued by the tracker Ethical AI Departures, which lists 68 profiles.
At Anthropic, the most prominent earlier exit was Mrinank Sharma, who had led the Safeguards Research Team since August 2023 and posted his resignation on 9 February, writing that “the world is in peril” and citing a gap between the company’s stated values and the decisions competitive pressure was driving.
Two other researchers, Behnam Neyshabur and Harsh Mehta, left in December 2025 to start a company rather than over any stated safety concern.
Neyshabur had co-led Anthropic’s Discovery team, which was building an “AI scientist”, and previously spent more than five years at Google DeepMind co-leading Gemini’s reasoning research; Mehta was a senior research scientist.
In June, the pair launched Mirendil, a startup building AI that automates AI research itself, and raised $200 million at a $1 billion valuation from Andreessen Horowitz, Kleiner Perkins and Nvidia—one of the largest seed rounds the sector has seen. The startup’s premise, that AI should improve itself autonomously to speed up science, sits at the centre of the recursive self-improvement debate that Coxon and Hubinger warn about.
At Google DeepMind, the trigger this year was the company’s April agreement letting the Pentagon use its Gemini models on classified networks. Alex Turner, who spent more than two years on the lab’s safety team, resigned in June, writing that “senior management had insisted that Google wouldn’t sign” and that once it did, “I couldn’t stay at Google in good conscience, so I left.”
He said the deal carried no binding restrictions against autonomous weapons or mass surveillance, and that he had declined outreach from OpenAI’s safety team and was working independently.
René Mayrhofer, a principal engineer for Android security, resigned over the same deal, writing that it left him “with the only choice to resign”. Another DeepMind scientist, Andreas Kirsch, called the contract’s safeguards “meaningless weasel words”. More than 580 Google staff, senior DeepMind researchers among them, signed a letter urging chief executive Sundar Pichai to reconsider; a group of London-based staff launched the first union bid at a frontier AI lab in May.
Google has defended the arrangement, saying its technology “is not intended for” and “should not be used for” domestic mass surveillance or autonomous weapons without human oversight—language which critics, including Turner, say carries no binding force. The company’s own chief scientist, Jeff Dean, has posted that mass surveillance “violates the Fourth Amendment,” and co-signed an amicus brief backing Anthropic against the Pentagon, but the deal proceeded regardless.
Anthropic was the outlier that refused to drop its restrictions on autonomous weapons and mass surveillance, a stance for which the Pentagon designated it a supply-chain risk.
OpenAI, too, has seen a spate of exits in recent months, though for distinct reasons. Zoë Hitzig resigned in February over the company’s move to place advertising inside ChatGPT, warning in an essay in The New York Times that it risked repeating social media’s error of optimising for engagement.
Ryan Beiermeister, the vice-president leading product policy, was fired in early January. OpenAI cited a discrimination allegation against her, which she called “absolutely false”. Beiermeister alleges the firing was retaliation for her opposition to a planned “adult mode” feature. OpenAI denies this, and the two accounts have yet to be reconciled.
Hieu Pham left after seven months citing burnout, writing “it’s when, not if” on the threat AI poses. Caitlin Kalinowski, the hardware and robotics lead, resigned over OpenAI’s Pentagon deal, writing that “surveillance of Americans without judicial oversight and lethal autonomy without human authorization are lines that deserved more deliberation than they got”.
The pattern predates this year. OpenAI dissolved its Superalignment team in 2024, prompting the exit of co-lead Jan Leike, who said “safety culture and processes have taken a backseat to shiny products,” and governance researcher Daniel Kokotajlo, who gave up about $1.7 million in vested equity rather than sign a non-disparagement agreement. Its co-founder and chief scientist Ilya Sutskever left the same year to start Safe Superintelligence Inc.
At xAI, most of the 11 original co-founders had departed by early 2026.
A wider warning
Coxon’s departure followed an essay, “An Alien Mind,” published on 6 September by OpenAI chief scientist Jakub Pachocki.
Pachocki wrote that “no lab has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer”.
He said he hoped “voluntary slowdowns” would “become commonplace until shared safety bars are established,” and called for international coordination to become “a top priority for governments around the world”.
Pachocki added that he was “concerned no one is prepared for the consequences of a continued rapid rise in machine intelligence”.
(Edited by Amrtansh Arora)
Also Read: Anthropic’s Claude Mythos: The AI model that India cannot access but cannot ignore either
