Cybersecurity Experts Say AI Labs Are Leaving Them Out of Safety Planning
The concern is not whether advanced models create cyber risk. It is whether the companies building safeguards have brought in the people needed to make those safeguards credible.
Listen to this story
The audio brief
Story brief
3 key pointsCybersecurity specialists are challenging how leading AI labs prepare for catastrophic AI-enabled hacking, arguing that published safeguards do not demonstrate sufficient input from practitioners who understand real attack methods. OpenAI says it paused two weeks of reinforcement-learning training, hardened research environments, expanded monitoring, and kept its largest planned frontier run on hold. Anthropic’s...
- 01
OpenAI paused two weeks of reinforcement-learning training while hardening environments and red-teaming them.
- 02
OpenAI’s controls include untrusted-code isolation, network separation, multistage monitoring, and automated investigation.
- 03
The company’s largest planned frontier reinforcement-learning run remained on hold during smaller training and evaluation work.
Cybersecurity experts told NBC News that leading AI companies are developing plans to prevent potentially catastrophic hacking campaigns without adequately involving security specialists or addressing fundamental cybersecurity problems. The criticism lands as the industry’s most prominent labs argue that more capable AI systems require stronger safeguards and, in some cases, a slower pace of development.
The gap identified by the experts is significant because the plans at issue are meant to address an extreme scenario: AI systems helping enable a hacking campaign with potentially catastrophic consequences. Their concern is not simply that companies should add another advisory panel. It is that basic cybersecurity questions are not being sufficiently addressed while labs devise their own safety approaches.
The safeguards labs say they are building
OpenAI has publicly described a broadening set of controls for advanced cyber-capable models. In an August post, the company said it temporarily slowed scaling and paused two weeks of reinforcement-learning training on its latest models intended for deployment while it hardened research environments, expanded monitoring, and red-teamed those environments. Its largest planned frontier reinforcement-learning run remained on hold while the company ran smaller training and evaluation work.
- Isolation requirements for workloads that execute untrusted or model-generated code.
- Network controls intended to keep higher-risk workloads separate from the internet and internal networks.
- Multistage monitoring designed to inspect concerning activity and escalate it to automated investigators.
Transparency is not the same as technical validation
Anthropic’s policy offers one route toward more outside scrutiny. Its Responsible Scaling Policy version 3.0 requires a published Frontier Safety Roadmap and says the company will appoint expert third-party reviewers in circumstances that require external review. The policy says those reviewers should know AI safety research, be incentivized to speak candidly, and be free of major conflicts of interest.
But the NBC-reported criticism points to a narrower test: whether outside review includes the right kind of expertise for the threat. A framework can require monitoring, access controls, model alignment, and public reporting, yet still leave unresolved how cybersecurity practitioners are involved in identifying weaknesses or judging whether defenses work against real attack methods.
A shared warning, but no shared answer
OpenAI CEO Sam Altman and Anthropic CEO Dario Amodei have both publicly supported slowing the race to build increasingly powerful AI systems. That agreement recognizes the pressure created by advancing capabilities. It does not answer the operational question raised by cybersecurity experts: who should help design and assess the defenses before the risk becomes harder to contain.
Sources
- openai.comPacing model development in an era of cyber-critical capabilities
- anthropic.comResponsible Scaling Policy Version 3.0
- nbcnews.comCybersecurity experts say AI giants are shutting them out of safety plans
Loading discussion...
Reader comments
Newest comments first. Replies stay oldest first.