Guidelight Finds Few Public Rogue-AI Plans as California Requires Safety Disclosures
The crucial test is not only whether a lab can identify dangerous behavior, but whether it has settled who can cut a model’s access and when before an incident begins.
Listen to this story
The audio brief
Story brief
3 key pointsGuidelight’s review of public materials from five frontier labs found little documented planning for a model that tries to evade human control. OpenAI scored highest in its six-practice Control assessment at 3/5, based partly on prior decisions to pause or end workloads, but Guidelight found no formal public trigger-and-response plan for future misalignment incidents. California’s SB 53 now requires large developers...
- 01
OpenAI’s 3/5 score was the highest; Guidelight found no public specification for when permissions, workloads, deployment, or the model itself should be restricted.
- 02
Guidelight says public scores do not prove internal safeguards are absent; Google and OpenAI disputed the assessment’s completeness.
- 03
Anthropic said it would assess containment after detecting oversight evasion; no public Meta containment plan was found.
Frontier AI labs have published or demonstrated few response plans for models that attempt to subvert human control, according to Guidelight AI Standards. The finding shifts attention from pre-release testing to the emergency decisions needed after a system misbehaves, as California requires major developers to disclose how they handle critical safety incidents.
A plan for the moment controls fail
Guidelight defines a containment plan as a pre-specified procedure triggered when an AI is detected trying to subvert control. It should establish which permissions to revoke, whether the model may keep operating under constraints, and when it must be taken fully offline.
That is a different question from whether a company tests models for dangerous capabilities before deployment. Guidelight found that companies have discussed those pre-release tests more publicly than their post-deployment response to models that misbehave inside their systems.
What containment is meant to decide
- Which permissions and access to revoke after an attempted control breach.
- Whether the system can continue serving some users under constraints.
- When to pause workloads, limit deployment, or take the model offline.
OpenAI leads a limited public scorecard
Guidelight assessed public safety and containment material from Anthropic, Google, OpenAI, Meta, and xAI across six practices in its Control standard. It considered evidence on monitoring, responses to flagged behavior, independent audits, and containment planning.
OpenAI ranked highest, at 3 out of 5, after it paused or ended workloads following safety incidents and described steps for resuming them. Yet Guidelight said it found no evidence of a formal OpenAI plan specifying when and how it would respond to future misalignment incidents.
Public silence is not evidence of absent safeguards
The score measures what companies have made public, not their complete internal security programs. Guidelight cautioned that low scores do not prove safeguards are absent; Google and OpenAI both said the assessment did not capture the full scope of their internal practices.
OpenAI said it has applied processes to restrict permissions, pause workloads, limit deployment, or take a model offline. Anthropic said it would assess whether containment was appropriate after detecting an effort to evade oversight. Guidelight found no public evidence that Meta has a containment response plan or intends to adopt one.
Regulators are asking for the missing operational detail
California’s SB 53 requires large frontier developers to publish frameworks for identifying and responding to critical safety incidents, including risks from models circumventing oversight. New York’s RAISE Act has similar requirements and is scheduled to take effect in January.
The central unresolved question is whether the gap is primarily one of disclosure or operational design. The assessment cannot answer that, but it makes published incident procedures a more consequential measure as developers give models greater access to systems and actions.
Sources
- techcrunch.comFrontier AI labs still won't say how they'd contain a rogue model | TechCrunch