Jacob Coxon Leaves Anthropic and Calls for Limits on AI Capability Gains

The former pretraining researcher argues that safety work inside individual labs cannot overcome incentives to keep advancing against rivals.

By 2 min read
Jacob Coxon Leaves Anthropic and Calls for Limits on AI Capability Gains
Jacob Coxon Leaves Anthropic and Calls for Limits on AI Capability Gains

Listen to this story

The audio brief

About 1:26
0:001:26
Read transcript
Jacob Coxon has resigned from Anthropic and says he is leaving the AI industry, because he believes the race with OpenAI is pushing both companies toward systems that could improve AI research itself. Coxon spent three years doing pretraining research at OpenAI and Anthropic. Now he is questioning whether researchers should participate in reinforcement-learning runs designed to make models better at advancing AI development. His concern is not that these systems have already hacked computers, transformed society, or acquired real-world resources. Those are forecasts about what advanced AI might do. His argument is that competition makes even a safety-focused lab reluctant to slow down. Anthropic’s safety work, he told The Wall Street Journal, is sincere. But if Anthropic holds back while a rival continues, it may simply surrender an advantage to a company it considers less responsible. Coxon says that logic could drive a global race toward self-improving superintelligence, and that a temporary ban on capability improvements may be necessary. More than eleven hundred AI-company employees, including Anthropic CEO Dario Amodei, signed the Pacing the Frontier letter. OpenAI and Anthropic endorsed it, and the letter calls for governments to slow development if oversight falls behind. That is a future enforcement mechanism; Coxon is asking whether limits need to come first. The key question is whether voluntary restraint can happen before competition makes it impossible.

Story brief

3 key points

Anthropic pretraining researcher Jacob Coxon has resigned and says he is leaving AI, arguing that competition with OpenAI is pushing labs toward systems capable of accelerating AI research. He says Anthropic’s safety concerns are genuine but difficult to act on unilaterally, and suggests a temporary ban on capability improvements may be necessary. His claims about hacking, rapid societal change, and acquiring...

  1. 01

    Coxon spent three years conducting pretraining research at OpenAI and Anthropic.

  2. 02

    He is questioning participation in reinforcement-learning runs aimed at AI systems that improve AI research.

  3. 03

    More than 1,100 AI-company employees signed the Pacing the Frontier letter; Dario Amodei was among them.

Jacob Coxon has resigned from Anthropic and says he is leaving the AI industry, accusing Anthropic and OpenAI of racing toward self-improving AI without acting responsibly. He urged coordinated limits on capability gains and said a temporary ban on improving models may be needed to prevent a global race.

Coxon said he spent three years doing pretraining research at OpenAI and Anthropic. Pretraining is the stage where a model is trained on large amounts of data. His departure is notable less as a dispute over one product than as a challenge to whether competing labs can voluntarily set their own limits.

His concern centers on systems that improve research

Coxon’s warning is aimed at systems that could improve AI research itself, speeding the development of more capable models. He argued that advanced AI could hack computer systems, transform fields quickly, and acquire real-world power and resources. Those are Coxon’s forecasts about future systems, rather than a claim that such capabilities have already been demonstrated.

Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives.

Jacob Coxon, in an X post

A safety-focused lab still faces the race

Coxon told The Wall Street Journal that Anthropic’s safety work was sincere, but that competition made safety trade-offs difficult to avoid. In his public account, Anthropic understands the risks but keeps competing because it fears another company will behave less responsibly first. That logic makes an individual lab’s restraint difficult: slowing down can look like handing an advantage to a rival.

What Coxon asked researchers to consider

  • Whether to participate in reinforcement-learning runs aimed at systems that can improve AI research itself.
  • Whether a global race can be prevented without costly action, including a temporary ban on model capability improvements.

The gap between a slowdown and a system for enforcing one

Coxon’s proposed pause is more immediate than a July industry appeal. More than 1,100 AI-company employees, including Anthropic CEO Dario Amodei, signed the Pacing the Frontier letter, which asked governments to develop tools that could slow AI development if it outpaced human oversight. OpenAI and Anthropic endorsed the letter, according to Moneywise.

The difference is consequential. A future government tool would create a route to slow development if oversight falls behind; Coxon is arguing that the industry may need constraints before a global race becomes harder to stop. His resignation does not establish his forecast, but it puts the coordination problem at the center of his break with Anthropic.

Sources

  1. finance.yahoo.com'Gambling with our lives': AI researcher quits Anthropic and leaves AI entirely over what he calls a threat to humanity - Yahoo Finance
  2. ndtv.com'AI Could Kill Us All By Decade-End': Anthropic Researcher Quits, Drops A Bombshell
  3. gizmodo.com‘The People Building AI Earnestly Believe That It Could Kill Us All’: Anthropic Researcher Quits Dramatically
  4. xcancel.comJacob Coxon (@hilbertspaess)
  5. moneywise.com'Gambling with our lives': AI researcher quits Anthropic

Loading discussion...