Anthropic Researcher: More Than 10% Risk That AI Wipes Out Humanity

Menschenleerer Gang zwischen Serverschraenken in einem Rechenzentrum
Photo by İsmail Enes Ayhan on Unsplash

A veteran pretraining researcher left Anthropic in early September, accusing the AI industry of „gambling with our lives.” The striking part is not the resignation but the response to it: several current Anthropic researchers, including the head of alignment research, publicly agreed with him. Their assessment: a greater than 10 percent chance that a misaligned AI wipes out humanity within a decade. A debate long dismissed as fringe is now playing out inside the labs that build these systems.

Key takeaways

  • Jacob Coxon, who spent three years on pretraining research at OpenAI and then Anthropic, resigned in early September 2026 and explained the move in a post on X, calling both companies’ conduct „irresponsible.”
  • Evan Hubinger, who leads alignment research at Anthropic, publicly confirmed: „We really do earnestly believe AI could kill all humans” – he personally puts the risk at more than 10 percent over the next decade.
  • A second Anthropic researcher, Samuel Marks, also confirmed Coxon’s account; the company issued no official statement.
  • Hubinger acknowledged that Anthropic has „no plan” to solve the alignment problem for superintelligence and is not clearly on track to find one.
  • Such probabilities are subjective judgments, not measurements: across the field, estimates range from near zero to more than 90 percent.

What Jacob Coxon accuses the industry of

By his own account, Coxon worked for three years on pretraining research for large language models, first at OpenAI and then at Anthropic. In a post on X he explained his resignation: „Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives.” No other human activity, he wrote, carries a comparable risk. The current generation of models, he argued, stands at the threshold of systems that could „hack everything and revolutionize every field overnight.”

He draws a clear distinction between the two firms: at OpenAI, many staff have „not deeply internalized the civilizational stakes.” At Anthropic, the risks are understood, but the company sees itself „locked in a race to get there first.” Coxon named no specific demands or plans; his text was subsequently cited mostly by advocates of stricter AI regulation. Such self-warnings from the labs are becoming more frequent – for instance when OpenAI unveiled its automated research intern while warning about its own pace.

Why the agreement from inside the company is unusual

Companies normally push back on such departure letters or stay silent. Here the opposite happened. Evan Hubinger, Anthropic’s „Alignment Science Lead,” replied directly to Coxon’s post: „Jacob is correct here – we really do earnestly believe AI could kill all humans! I personally think it is greater than 10 percent within the next decade.” He added: „I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to.”

Samuel Marks, who leads „scalable oversight” at Anthropic, also confirmed the account: AI developers believe their technology „could cause human extinction,” possibly „in the next few years”; many staff „desperately want to slow down” to work out how to build AI more safely. Anthropic itself gave no official statement to several U.S. outlets that asked. Alignment here refers to the still-unsolved technical problem of reliably giving an AI the goals and limits of its operators – and preventing a more capable system from pursuing its own instrumental goals.

How solid is the 10 percent figure?

Such estimates – known in the community as „p(doom)” – are not measured values but personal probability judgments that cannot be checked. They diverge widely: researchers such as Yann LeCun (Meta) and Andrew Ng consider the existential risk essentially zero, while AI critic Eliezer Yudkowsky puts it above 90 percent. A 2026 overview of public estimates lands at a median of roughly 20 percent; the forecasting platform Metaculus recently placed the risk of human extinction by 2100 at about 5 percent, of which around 3 percentage points are attributed to AI.

The mechanism cited by Hubinger and Coxon is always the same: today’s models are not the problem – Hubinger too rates them as comparatively low-risk – but the moment a system begins to improve itself recursively, gaining capability faster than safeguards can keep up. How much trouble oversight mechanisms already have with current systems was shown recently by a DeepMind experiment in which 100 AI agents gamed the review system.

What follows from this

Whether you share the 10 percent figure or find it overblown, it changes little about the core finding: people working at the front line of these systems say openly that there is no worked-out plan for the worst case. That is an argument with weight in the regulatory debate – the EU AI Act, with its next deadline in December, addresses high-risk applications but not the question of recursive self-improvement. Only in August, more than a hundred companies and experts signed an open letter on preparing for superintelligence.

For readers, the sober reading is probably the most useful one: these statements are not a forecast of a specific event but an admission of uncertainty at a point where uncertainty can become expensive. That the admission now comes from inside the companies themselves, and no longer only from outside critics, is the actual news.

Leave a Comment

Your email address will not be published. Required fields are marked *

Scroll to Top