Cybersecurity Experts Say AI Labs Are Leaving Them Out of Safety Planning

The concern is not whether advanced models create cyber risk. It is whether the companies building safeguards have brought in the people needed to make those safeguards credible.

By 2 min read
Cybersecurity Experts Say AI Labs Are Leaving Them Out of Safety Planning
Cybersecurity Experts Say AI Labs Are Leaving Them Out of Safety Planning

Listen to this story

The audio brief

About 1:35
0:001:35
Read transcript
OpenAI paused two weeks of reinforcement-learning training while it hardened its research environments, expanded monitoring, and red-teamed those systems. The company says its safeguards now include isolating workloads that run untrusted or model-generated code, separating higher-risk systems from the internet and internal networks, and using several layers of monitoring to flag suspicious activity for automated investigation. Its largest planned frontier training run was still on hold while smaller training and evaluation work continued. That sounds like a serious response to serious risk. But cybersecurity experts told NBC News that the deeper problem is who is helping design and test these defenses. The concern is not only whether advanced AI could enable a catastrophic hacking campaign. It is whether the labs’ safety plans are being shaped by practitioners who understand how real attacks work, and whether the controls have been technically validated against those methods. Anthropic’s Responsible Scaling Policy, version three, requires a Frontier Safety Roadmap and, in some cases, review by outside experts. The policy says those reviewers should understand AI safety research, speak candidly, and avoid major conflicts of interest. But transparency and external review do not necessarily guarantee meaningful cybersecurity scrutiny. OpenAI CEO Sam Altman and Anthropic CEO Dario Amodei have both supported slowing the race toward more capable systems. The unresolved question is more operational: are the right security specialists involved before these safeguards are asked to contain risks that may be much harder to control?

Story brief

3 key points

Cybersecurity specialists are challenging how leading AI labs prepare for catastrophic AI-enabled hacking, arguing that published safeguards do not demonstrate sufficient input from practitioners who understand real attack methods. OpenAI says it paused two weeks of reinforcement-learning training, hardened research environments, expanded monitoring, and kept its largest planned frontier run on hold. Anthropic’s...

  1. 01

    OpenAI paused two weeks of reinforcement-learning training while hardening environments and red-teaming them.

  2. 02

    OpenAI’s controls include untrusted-code isolation, network separation, multistage monitoring, and automated investigation.

  3. 03

    The company’s largest planned frontier reinforcement-learning run remained on hold during smaller training and evaluation work.

Cybersecurity experts told NBC News that leading AI companies are developing plans to prevent potentially catastrophic hacking campaigns without adequately involving security specialists or addressing fundamental cybersecurity problems. The criticism lands as the industry’s most prominent labs argue that more capable AI systems require stronger safeguards and, in some cases, a slower pace of development.

The gap identified by the experts is significant because the plans at issue are meant to address an extreme scenario: AI systems helping enable a hacking campaign with potentially catastrophic consequences. Their concern is not simply that companies should add another advisory panel. It is that basic cybersecurity questions are not being sufficiently addressed while labs devise their own safety approaches.

The safeguards labs say they are building

OpenAI has publicly described a broadening set of controls for advanced cyber-capable models. In an August post, the company said it temporarily slowed scaling and paused two weeks of reinforcement-learning training on its latest models intended for deployment while it hardened research environments, expanded monitoring, and red-teamed those environments. Its largest planned frontier reinforcement-learning run remained on hold while the company ran smaller training and evaluation work.

  • Isolation requirements for workloads that execute untrusted or model-generated code.
  • Network controls intended to keep higher-risk workloads separate from the internet and internal networks.
  • Multistage monitoring designed to inspect concerning activity and escalate it to automated investigators.

Transparency is not the same as technical validation

Anthropic’s policy offers one route toward more outside scrutiny. Its Responsible Scaling Policy version 3.0 requires a published Frontier Safety Roadmap and says the company will appoint expert third-party reviewers in circumstances that require external review. The policy says those reviewers should know AI safety research, be incentivized to speak candidly, and be free of major conflicts of interest.

But the NBC-reported criticism points to a narrower test: whether outside review includes the right kind of expertise for the threat. A framework can require monitoring, access controls, model alignment, and public reporting, yet still leave unresolved how cybersecurity practitioners are involved in identifying weaknesses or judging whether defenses work against real attack methods.

A shared warning, but no shared answer

OpenAI CEO Sam Altman and Anthropic CEO Dario Amodei have both publicly supported slowing the race to build increasingly powerful AI systems. That agreement recognizes the pressure created by advancing capabilities. It does not answer the operational question raised by cybersecurity experts: who should help design and assess the defenses before the risk becomes harder to contain.

Sources

  1. openai.comPacing model development in an era of cyber-critical capabilities
  2. anthropic.comResponsible Scaling Policy Version 3.0
  3. nbcnews.comCybersecurity experts say AI giants are shutting them out of safety plans

Loading discussion...