Anthropic Proposes Embedded AI Safety Evaluators Without Power to Stop Releases
Dario Amodei’s plan would give third parties unusual access and publication rights. Its banking comparison also shows the distance between scrutiny and enforceable oversight.
Listen to this story
The audio brief
Story brief
3 key pointsAnthropic CEO Dario Amodei is proposing a permanent, independent evaluation presence inside frontier AI labs, with access similar to internal safety teams and the right to publish findings after limited redactions. The major constraint is enforcement: evaluators could investigate and report risks but could not block training or releases. METR is a named candidate, while the proposal leaves unresolved who selects...
- 01
Evaluators would inspect models, surrounding instructions, tools, permissions, safety controls, and action logs—not just prompt outputs.
- 02
The banking analogy is limited: bank examiners can compel corrective action; Anthropic’s reviewers could not.
- 03
XBOW’s Albert Ziegler warns testing may miss rare combinations that produce catastrophic outcomes.
Anthropic CEO Dario Amodei has proposed placing third-party safety evaluators inside frontier AI companies on an ongoing basis. The evaluators would receive access comparable to internal risk teams and could publish findings with limited redactions, but they would not have legal authority to stop a model from being trained or released.
Amodei’s proposal would make external evaluation a continuing presence at companies building the most capable AI systems, rather than a one-time review. He pointed to bank supervision as a precedent: government examiners can work inside banks with access to systems and employees.
What Anthropic says evaluators would get
- Access comparable to Anthropic’s internal risk-assessment teams.
- The ability to investigate systems and report findings.
- Publication rights without Anthropic’s editorial control, subject to limited redactions.
Albert Ziegler, head of AI at cybersecurity company XBOW, said evaluations can expose important model failures but may not trigger the rare mix of conditions that could produce a catastrophic outcome. He described that as a limit of testing, not evidence that such an outcome has been observed.
Assessing a whole AI system may also require more than prompting a model. Ziegler said reviewers may need to inspect the instructions surrounding it, its tools and permissions, safety controls, and logs of attempted actions. A serious problem could still emerge in circumstances a review never reaches.
Critics cited by CNBC argue that independence can be hard to establish when a developer selects the evaluator, sets the bounds of access, and decides what to do after a finding. In that arrangement, they say, an outside reviewer can look more like an internal compliance function than a regulator.
Amodei named the nonprofit Model Evaluation and Threat Research, or METR, as one possible embedded evaluator. Anthropic has worked with METR on evaluations before, offering a concrete example of the organizations that could take on the role.
The proposal would create a route for outsiders to see more of a frontier lab’s safety work and speak publicly about it. Whether that transparency changes a company’s choices will depend on how evaluators are selected, what they can examine, and how companies respond when a review identifies a risk.
Sources
- cnbc.comAnthropic, OpenAI's proposed AI risk evaluators may not have enough power to prevent disasters
Loading discussion...
Reader comments
Newest comments first. Replies stay oldest first.