Evan Hubinger Backs Coxon’s AI Warning With a 10%+ Extinction Estimate

The public response from Anthropic’s alignment leader adds a personal probability estimate to a viral warning—and sharpens the question of how labs would slow down.

By 2 min read
Evan Hubinger Backs Coxon’s AI Warning With a 10%+ Extinction Estimate
Evan Hubinger Backs Coxon’s AI Warning With a 10%+ Extinction Estimate

Listen to this story

The audio brief

About 1:32
0:001:32
Read transcript
Evan Hubinger, an alignment leader at Anthropic, has publicly endorsed Jacob Coxon’s warning about the AI race—and put a stark number on the risk. Hubinger says his personal forecast gives AI a greater-than-ten-percent chance of killing all humans within the next decade. He also says Anthropic does not yet have a clear plan for aligning superintelligence with human goals. That is a personal estimate, not Anthropic’s official risk assessment, and it is not evidence that superintelligence exists today. The intervention follows Coxon’s resignation from Anthropic. Coxon, who previously worked on pretraining at OpenAI, accused both companies of racing toward self-improving superintelligence and gambling with human lives. His post on X drew more than 156 million views and 752,000 likes, while Illinois Governor JB Pritzker called for immediate action. The dispute then reached CNN and Fox News, pushing an internal safety argument into mainstream politics. The practical issue is what happens next. Coxon has proposed that OpenAI and Anthropic limit recursive self-improvement—using AI to build better AI. Anthropic told WIRED it supports a lawful, verifiable way to pace powerful-model releases, but there is no operating agreement with OpenAI. The key test is whether those broad principles become defined, checkable release limits before more capable systems are deployed.

Story brief

3 key points

Anthropic alignment researcher Evan Hubinger has endorsed former colleague Jacob Coxon’s warning that competitive development could outrun safety controls, while putting a personal number on the downside: more than 10% odds AI kills everyone within the next decade. Hubinger also said Anthropic lacks a clear plan for superintelligence alignment. The immediate market and policy question is not whether extinction is...

  1. 01

    Hubinger’s figure is a personal forecast, not Anthropic’s official risk estimate or evidence that superintelligence exists today.

  2. 02

    Coxon’s X thread drew 156M+ views and 752,000+ likes; Illinois Gov. JB Pritzker called for immediate action.

  3. 03

    Anthropic told WIRED it supports lawful, verifiable pacing, but no operating agreement with OpenAI exists.

Anthropic alignment-stress-testing leader Evan Hubinger has publicly backed former colleague Jacob Coxon’s warning about the AI race. Hubinger said he personally sees a greater than 10% chance that AI could kill all humans within the next decade, and that Anthropic does not yet have a plan to solve superintelligence alignment.

Hubinger’s response came after Coxon resigned from Anthropic, where he had worked on pretraining after time at OpenAI. Coxon accused both companies of racing toward self-improving superintelligence and gambling with human lives. That is Coxon’s assessment of the companies’ direction, not evidence that such a system exists today.

Hubinger’s personal assessment
Greater than 10%Chance AI kills all humans within a decade

Hubinger said Coxon was correct, while stressing that his probability estimate was his personal forecast about a future outcome.

The warning reached beyond AI circles

CNET counted more than 156 million views and more than 752,000 likes on Coxon’s X thread at the time of publication. Illinois Gov. JB Pritzker responded by calling for immediate action from the AI industry and Washington. Coxon’s warning also reached CNN and Fox News, a sign that an internal AI-safety dispute had moved into mainstream political conversation.

A risk estimate without a settled control plan

Alignment is the effort to make an AI system reliably follow intended human goals. Hubinger said Anthropic is trying its best, but is not clearly on track to solve alignment for superintelligence. His statement puts a senior member of the company’s safety work alongside Coxon’s broader concern: that competitive pressure may outrun the ability to control more capable systems.

The next test is a verifiable limit

Coxon has proposed that OpenAI and Anthropic coordinate to limit recursive self-improvement, or using AI to create better AI systems. Anthropic told WIRED it supports a lawful, verifiable way for the industry to pace releases of powerful models. Neither is an operating agreement. The concrete test for labs and policymakers is whether those broad calls become defined limits that can be checked before more capable systems are released.

The full extinction scenario remains an extreme and contested forecast, and it is unclear whether current technology can produce superintelligence. But Hubinger’s intervention changes the immediate debate: it is no longer only about a departing researcher’s warning, but about whether leading labs can specify the safety conditions they say must govern the race.

Sources

  1. wired.comThe AI Researcher Who Just Quit Anthropic Says It’s ‘Crunch Time for Humanity’
  2. cnet.comWhat an Ex-Anthropic Researcher's Warning About Human Extinction Really Means - CNET
  3. the-decoder.comAI safety panic goes mainstream after Anthropic researcher's warnings land on CNN and Fox News

Loading discussion...