Stuart Russell Says AI Progress Should Wait for Safety Proof
The AI researcher’s new critique of Dario Amodei’s pacing proposal turns on a consequential distinction: extra time for safety research versus rules that can stop development altogether.
Listen to this story
The audio brief
Story brief
3 key pointsThe fight over frontier-AI safety is becoming a sequencing dispute. Anthropic CEO Dario Amodei’s “Pace the Frontier” plan would slow capability gains while research continues, potentially creating one to two years for interpretability, alignment, and testing. Stuart Russell argues that schedule-based restraint is insufficient: developers should advance only after models pass capability-linked safety tests and...
- 01
Amodei’s proposal aims to buy one to two years for interpretability, alignment, and testing work.
- 02
Russell compares AI advancement to aircraft certification: passing predefined tests must precede service, not follow a calendar.
- 03
The framework proposes embedding independent evaluators inside frontier companies with full system access.
Stuart Russell is challenging the idea that frontier AI can be made safer simply by moving more slowly. In a newly published article, the UC Berkeley computer scientist argues that developers should advance increasingly capable systems only after they meet predefined safety requirements and demonstrate specified alignment properties. The distinction could determine whether a proposed slowdown is a delay—or a genuine stop sign.
Russell’s target is Dario Amodei’s “Pace the Frontier” proposal. Anthropic’s chief executive has called for a framework intended to reduce risks from rapidly advancing AI while continuing model training and technical progress. Amodei has said pacing does not mean stopping either activity; he argues that a slower rate of capability improvement could create one to two additional years for work on interpretability, alignment, and testing.
What Amodei’s proposal asks companies to do
- Put third-party AI-system evaluators inside companies with full access to their systems.
- Have frontier AI companies in democratic countries develop common safety standards and limits on unchecked progress, with regulation where needed.
- Extend the effort to authoritarian countries through a broader compact.
The proposal has attracted support, according to Russell’s article, from OpenAI’s Sam Altman, xAI’s Elon Musk, Google DeepMind’s Demis Hassabis, and Microsoft’s Satya Nadella. That backing gives the argument over its design more weight than a disagreement between one company and one critic.
A disagreement about sequencing
The disagreement is not over whether safety work is needed. It is over what comes first. Russell rejects setting a slower schedule for capability gains and hoping researchers can use the resulting time well. He instead argues that safety requirements are non-negotiable: further progress should occur only when developers can show they have met them.
Russell uses aircraft certification as the model. A manufacturer is not permitted to introduce a new plane on a fixed annual schedule and then hope testing works out, he writes. The aircraft enters service only after passing tests and receiving certification. For AI, his equivalent is a rule that a model with a particular capability must come with certifications of particular alignment properties.
When a slowdown becomes a halt
That ordering produces a much tougher outcome than a voluntary speed limit. Russell describes it as consistent with the “red lines” approach advocated by AI safety researchers. If developers cannot demonstrate that a more capable system is safe enough to deploy, he argues, progress may have to stop. In his framing, the industry would face a red flag rather than a pace car.
Russell also sees a partial opening in Amodei’s own language. Amodei has proposed rules stating that if models possess a given capability, they should be accompanied by certifications of specified alignment properties. Russell says that formulation is closer to a requirements-first system than the pacing metaphor suggests. The unresolved issue is whether leading AI companies and governments would accept standards stringent enough to block further progress when the evidence falls short.
Sources
- theguardian.comAI safety requires more than just slowing our pace | Stuart Russell
Loading discussion...
Reader comments
Newest comments first. Replies stay oldest first.