Stuart Russell Says AI Progress Should Wait for Safety Proof

The AI researcher’s new critique of Dario Amodei’s pacing proposal turns on a consequential distinction: extra time for safety research versus rules that can stop development altogether.

By 3 min read
Stuart Russell Says AI Progress Should Wait for Safety Proof
Stuart Russell Says AI Progress Should Wait for Safety Proof

Listen to this story

The audio brief

About 1:39
0:001:39
Read transcript
Stuart Russell is challenging a central idea in frontier-AI safety: that moving slowly is enough. In a newly published article, the computer scientist argues that increasingly capable models should advance only after they pass predefined safety tests and demonstrate specific alignment properties. That is a much stronger condition than simply adding time to the development schedule. The target is Dario Amodei’s “Pace the Frontier” proposal. Amodei, Anthropic’s chief executive, says slower capability gains could create one to two extra years for interpretability, alignment, and testing research—without stopping model training or technical progress. His framework would also place independent evaluators inside frontier companies, with full access to their systems, and encourage democratic countries to establish shared safety standards and limits on unchecked progress. The proposal reportedly has support from OpenAI’s Sam Altman, xAI’s Elon Musk, Google DeepMind’s Demis Hassabis, and Microsoft’s Satya Nadella. So this is not just a dispute between one company and one critic. It is an argument about sequencing: should safety research happen during a managed slowdown, or should capability gains wait until safety evidence is in? Russell compares AI to aircraft certification. A new plane does not enter service because the calendar says it is time; it flies after passing required tests. Applied to AI, that could mean a slowdown becomes a halt whenever a more capable system cannot prove it is safe enough to deploy. The question now is whether companies and governments would accept standards strict enough to stop progress when the evidence falls short.

Story brief

3 key points

The fight over frontier-AI safety is becoming a sequencing dispute. Anthropic CEO Dario Amodei’s “Pace the Frontier” plan would slow capability gains while research continues, potentially creating one to two years for interpretability, alignment, and testing. Stuart Russell argues that schedule-based restraint is insufficient: developers should advance only after models pass capability-linked safety tests and...

  1. 01

    Amodei’s proposal aims to buy one to two years for interpretability, alignment, and testing work.

  2. 02

    Russell compares AI advancement to aircraft certification: passing predefined tests must precede service, not follow a calendar.

  3. 03

    The framework proposes embedding independent evaluators inside frontier companies with full system access.

Stuart Russell is challenging the idea that frontier AI can be made safer simply by moving more slowly. In a newly published article, the UC Berkeley computer scientist argues that developers should advance increasingly capable systems only after they meet predefined safety requirements and demonstrate specified alignment properties. The distinction could determine whether a proposed slowdown is a delay—or a genuine stop sign.

Russell’s target is Dario Amodei’s “Pace the Frontier” proposal. Anthropic’s chief executive has called for a framework intended to reduce risks from rapidly advancing AI while continuing model training and technical progress. Amodei has said pacing does not mean stopping either activity; he argues that a slower rate of capability improvement could create one to two additional years for work on interpretability, alignment, and testing.

What Amodei’s proposal asks companies to do

  • Put third-party AI-system evaluators inside companies with full access to their systems.
  • Have frontier AI companies in democratic countries develop common safety standards and limits on unchecked progress, with regulation where needed.
  • Extend the effort to authoritarian countries through a broader compact.

The proposal has attracted support, according to Russell’s article, from OpenAI’s Sam Altman, xAI’s Elon Musk, Google DeepMind’s Demis Hassabis, and Microsoft’s Satya Nadella. That backing gives the argument over its design more weight than a disagreement between one company and one critic.

A disagreement about sequencing

The disagreement is not over whether safety work is needed. It is over what comes first. Russell rejects setting a slower schedule for capability gains and hoping researchers can use the resulting time well. He instead argues that safety requirements are non-negotiable: further progress should occur only when developers can show they have met them.

Russell uses aircraft certification as the model. A manufacturer is not permitted to introduce a new plane on a fixed annual schedule and then hope testing works out, he writes. The aircraft enters service only after passing tests and receiving certification. For AI, his equivalent is a rule that a model with a particular capability must come with certifications of particular alignment properties.

When a slowdown becomes a halt

That ordering produces a much tougher outcome than a voluntary speed limit. Russell describes it as consistent with the “red lines” approach advocated by AI safety researchers. If developers cannot demonstrate that a more capable system is safe enough to deploy, he argues, progress may have to stop. In his framing, the industry would face a red flag rather than a pace car.

Russell also sees a partial opening in Amodei’s own language. Amodei has proposed rules stating that if models possess a given capability, they should be accompanied by certifications of specified alignment properties. Russell says that formulation is closer to a requirements-first system than the pacing metaphor suggests. The unresolved issue is whether leading AI companies and governments would accept standards stringent enough to block further progress when the evidence falls short.

Sources

  1. theguardian.comAI safety requires more than just slowing our pace | Stuart Russell

Loading discussion...