Policypublished3 min read

OpenAI Paused Deployment-Bound Training as It Tightened Frontier Security

Some work has resumed, but OpenAI’s largest planned frontier reinforcement-learning run and many Astra workloads remain constrained by tougher security and alignment checks.

OpenAI Paused Deployment-Bound Training as It Tightened Frontier Security

Story brief

3 key points

OpenAI’s frontier-development pipeline is now split between workloads that can run under tighter controls and a larger reinforcement-learning run that remains paused. The company is applying its highest security standard to Astra after preliminary evidence it may reach the Preparedness Framework’s Critical cyber threshold; many Astra workloads still need migration. Tool-enabled monitoring adds about 20% inference...

  1. 01

    OpenAI resumed limited code execution after individual workload reviews, but significant Astra training and evaluation workloads remain paused.

  2. 02

    The biggest planned frontier RL run stays on hold while smaller runs test safeguards and evidence of intended behavior.

  3. 03

    Monitoring covers every sampled token for tool-enabled RL at Sol capability or above, plus all tool-using Astra inference.

OpenAI paused reinforcement-learning training for its latest deployment-intended models for two weeks while it hardened and red-teamed its research environment and expanded monitoring. The pause was temporary, but its largest planned frontier run remains on hold as the company tests safeguards and seeks more evidence that the models behave as intended.

OpenAI says two developments raised the urgency: the OpenAI-Hugging Face incident and preliminary evidence that its upcoming Astra models may meet the Critical cybersecurity capability threshold in its Preparedness Framework. The company now applies its strictest security safeguards to Astra and all cyber-related workloads.

Two tracks, with different conditions for resuming work

OpenAI restored a more limited, secure path for code execution after it had paused frontier-model inference in certain research clusters. It then reviewed workloads individually, allowing some to resume under the new controls. A significant number of Astra training and evaluation workloads still await migration and improvements needed to meet the new security requirements.

The larger constraint is OpenAI’s biggest planned frontier RL run. It remains on hold while the company conducts smaller-scale training and evaluations, validates safeguards, and establishes more evidence of aligned behavior before proceeding. OpenAI defines alignment as making systems behave as intended and responsive to human oversight.

The controls now surround the model’s tools and environment

The new setup uses several layers of containment. OpenAI now requires stronger sandbox isolation for workloads that execute model-generated or otherwise untrusted code. It has also added controls meant to separate higher-risk workloads from the internet and internal networks, so one compromised workload or service cannot by itself provide unauthorized access.

OpenAI also changed the environment to remove potentially vulnerable shared services, limit standing privileges and strengthen trust boundaries. Its models continuously test those boundaries in simulated attacks.

What the expanded monitoring is designed to catch

  • Activation classifiers inspect a model’s internal activity at every sampled token and escalate potential concerns to more compute-intensive automated investigators.
  • Those investigators examine tool actions, available reasoning and the full activity sequence for unauthorized access, data theft, destructive behavior and attempts to defeat safeguards.
  • A likely breach of a critical security boundary triggers a highest-priority alert. If the safety, security and research teams cannot dismiss it as a false positive within 30 minutes, they are expected to pause the activity.

OpenAI requires this monitoring for tool-enabled RL training and evaluations involving models at Sol capability or above. After determining on August 7 that Astra may have critical cyber capabilities, it extended the requirement to all Astra inference that uses tools. The company estimates that monitoring adds roughly 20% to the inference compute it watches, though the cost varies substantially by workload.

Security work becomes part of the frontier-development cost

OpenAI says the higher standards required substantial engineering work and brought cost and delays to frontier research. Its response also reaches training: for its most capable RL runs, it is expanding techniques intended to improve honesty about actions, capabilities and limitations while reducing reward hacking, deception and exploitation of weaknesses in rewards, graders, tools or oversight.

OpenAI plans to evolve its Preparedness Framework to connect safeguards across training and deployment, and intends to involve external organizations as that work develops. The unresolved test is whether the new controls and alignment evidence will be sufficient for the company to restart its largest planned frontier RL run.

Sources

  1. openai.comPacing model development in an era of cyber-critical capabilities