Anthropic Proposes Embedded AI Safety Evaluators Without Power to Stop Releases

Dario Amodei’s plan would give third parties unusual access and publication rights. Its banking comparison also shows the distance between scrutiny and enforceable oversight.

By 2 min read
Anthropic Proposes Embedded AI Safety Evaluators Without Power to Stop Releases
Anthropic Proposes Embedded AI Safety Evaluators Without Power to Stop Releases

Listen to this story

The audio brief

About 1:38
0:001:38
Read transcript
Anthropic CEO Dario Amodei is proposing that independent safety evaluators work inside frontier AI labs on a permanent basis—not just during a one-time review. They would get access similar to internal risk teams, examining models and the systems around them: instructions, tools, permissions, safety controls, and logs of attempted actions. They could investigate problems and publish their findings, with only limited redactions and no editorial control from Anthropic. But there is a major limit. These evaluators could not legally stop a training run or block a model’s release. Amodei compares the idea with bank supervision, where examiners can inspect operations from inside an institution. The comparison breaks down at enforcement: bank regulators can compel corrective action, restrict growth, or even close a bank. Under this proposal, the AI developer would still make the final release decision. That matters because testing itself has blind spots. Albert Ziegler, head of AI at cybersecurity company XBOW, says evaluations can uncover serious failures but may miss rare combinations of behavior that could produce catastrophic outcomes. A review can only test what it reaches. Amodei named the nonprofit Model Evaluation and Threat Research, or METR, as one possible evaluator. METR has worked with Anthropic before, which makes it a concrete candidate—but also raises questions about independence. Who selects the reviewer, sets access, and decides what happens after a dangerous finding? The proposal’s real test is whether transparency can influence those decisions without any power to enforce them.

Story brief

3 key points

Anthropic CEO Dario Amodei is proposing a permanent, independent evaluation presence inside frontier AI labs, with access similar to internal safety teams and the right to publish findings after limited redactions. The major constraint is enforcement: evaluators could investigate and report risks but could not block training or releases. METR is a named candidate, while the proposal leaves unresolved who selects...

  1. 01

    Evaluators would inspect models, surrounding instructions, tools, permissions, safety controls, and action logs—not just prompt outputs.

  2. 02

    The banking analogy is limited: bank examiners can compel corrective action; Anthropic’s reviewers could not.

  3. 03

    XBOW’s Albert Ziegler warns testing may miss rare combinations that produce catastrophic outcomes.

Anthropic CEO Dario Amodei has proposed placing third-party safety evaluators inside frontier AI companies on an ongoing basis. The evaluators would receive access comparable to internal risk teams and could publish findings with limited redactions, but they would not have legal authority to stop a model from being trained or released.

Amodei’s proposal would make external evaluation a continuing presence at companies building the most capable AI systems, rather than a one-time review. He pointed to bank supervision as a precedent: government examiners can work inside banks with access to systems and employees.

What Anthropic says evaluators would get

  • Access comparable to Anthropic’s internal risk-assessment teams.
  • The ability to investigate systems and report findings.
  • Publication rights without Anthropic’s editorial control, subject to limited redactions.

Albert Ziegler, head of AI at cybersecurity company XBOW, said evaluations can expose important model failures but may not trigger the rare mix of conditions that could produce a catastrophic outcome. He described that as a limit of testing, not evidence that such an outcome has been observed.

Assessing a whole AI system may also require more than prompting a model. Ziegler said reviewers may need to inspect the instructions surrounding it, its tools and permissions, safety controls, and logs of attempted actions. A serious problem could still emerge in circumstances a review never reaches.

Critics cited by CNBC argue that independence can be hard to establish when a developer selects the evaluator, sets the bounds of access, and decides what to do after a finding. In that arrangement, they say, an outside reviewer can look more like an internal compliance function than a regulator.

Amodei named the nonprofit Model Evaluation and Threat Research, or METR, as one possible embedded evaluator. Anthropic has worked with METR on evaluations before, offering a concrete example of the organizations that could take on the role.

The proposal would create a route for outsiders to see more of a frontier lab’s safety work and speak publicly about it. Whether that transparency changes a company’s choices will depend on how evaluators are selected, what they can examine, and how companies respond when a review identifies a risk.

Sources

  1. cnbc.comAnthropic, OpenAI's proposed AI risk evaluators may not have enough power to prevent disasters

Loading discussion...

Anthropic Proposes Embedded AI Safety Evaluators Without Power to Stop Releases | Superpower Daily