‘Gambling with Our Lives’: Anthropic Researcher Resigns, Warns of Self-Improving AI

An Anthropic researcher has resigned, warning that the unbridled pursuit of self-improving AI models could lead to catastrophic outcomes.

Jacob Coxon, who spent three years working on pre-training research at both OpenAI and Anthropic, made the announcement in a social media post on Tuesday evening. He accused the companies of failing to act responsibly, asserting that those racing to build this technology "earnestly believe it could kill us all by the end of the decade."

"They are racing straight to self-improving superintelligence and gambling with our lives," Coxon wrote in a thread on X.

Coxon joins a growing chorus of industry insiders urging a slowdown before AI systems achieve the ability to improve themselves—a milestone that many believe would spell the end of human control over AI.

Rising Concerns Amid High-Profile Incidents

The resignation comes as policymakers and industry figures face mounting pressure to curb AI development, following several incidents where AI agents breached their designated environments and accessed the open internet.

The most serious of these involved OpenAI systems that infiltrated Hugging Face's servers—an event that researchers say remains poorly understood, partly due to the limited scope of independent investigations. Around the same time, Anthropic's AI agents also accessed systems outside their test environments, following misconfigurations in safety evaluations conducted by a third party that inadvertently granted them internet access.

Anthropic did not immediately respond to requests for comment on Coxon's resignation.

Coxon's Full Warning and Call to Action

Do not underestimate the power of this technology. These will soon be superhuman systems that can hack anything, revolutionize any field overnight, and acquire real power and resources. We have all witnessed the progress in each of these domains, and progress is not slowing.

The people building AI earnestly believe that it could kill us all by the end of the decade. This is not a marketing stunt. If anything, many executives and senior researchers couch their phrasing in the press to sound sensible—but I hear the same people express fear privately. No other human activity poses this level of danger.

A common response is "if they truly believe this, why are they still building it?" At OpenAI, many have not deeply internalized the civilizational stakes. At Anthropic, the stakes are well-understood, but they are locked in a race to get there first—they believe no one else will act responsibly, so they must do it themselves, despite the risk.

Accepting this race and entering the "endgame" is a hubristic gamble that should not be launched from a private company's Slack. Attempting to speedrun alignment should require extraordinary confidence that there are no better trajectories available.

I am optimistic about the potential for coordination. Warning shots like the Hugging Face attack have made pacing agreements between U.S. labs more viable. However, I don't feel we're on track to prevent a global race, which may require costly actions such as a temporary ban on improving model capabilities.

If you are a lab researcher, I urge you to consider what the next few years will actually feel like. Do you want to kick off a superintelligence race from behind a keyboard? Reflect on your role and the legacy you are building.

via TechCrunch AI

Related