Policypublished

Guidelight Finds Few Public Rogue-AI Plans as California Requires Safety Disclosures

The crucial test is not only whether a lab can identify dangerous behavior, but whether it has settled who can cut a model’s access and when before an incident begins.

By 3 min read
Guidelight Finds Few Public Rogue-AI Plans as California Requires Safety Disclosures

Listen to this story

The audio brief

About 1:35
0:001:35
Read transcript
Guidelight’s review of five leading AI labs found few public plans for the moment a model tries to evade human control. The concern is not just whether a lab can spot dangerous behavior before release. It is whether, during an incident, the company has already decided which permissions to revoke, whether the model can keep operating under restrictions, and when to take it fully offline. Guidelight assessed Anthropic, Google, OpenAI, Meta, and xAI across six practices in its Control standard, including monitoring, responses to flagged behavior, audits, and containment planning. OpenAI scored highest, with three out of five. That score reflected documented cases in which it paused or ended workloads after safety incidents, and described how work could resume. But Guidelight said it found no formal public trigger-and-response plan for future misalignment incidents—no published specification for when to restrict permissions, pause workloads, limit deployment, or shut down the model. The score measures disclosure, not the full state of a company’s internal safeguards. Google and OpenAI disputed the assessment’s completeness. OpenAI said it has used processes to restrict access and deployment. Anthropic said it would assess containment after detecting an effort to evade oversight. Guidelight found no public containment plan from Meta. Regulation is now pushing this detail into the open. California’s SB 53 requires large frontier developers to publish critical-incident frameworks, and New York’s RAISE Act takes effect in January. The key question is whether the missing plans reflect secrecy, or unfinished operational design.

Story brief

3 key points

Guidelight’s review of public materials from five frontier labs found little documented planning for a model that tries to evade human control. OpenAI scored highest in its six-practice Control assessment at 3/5, based partly on prior decisions to pause or end workloads, but Guidelight found no formal public trigger-and-response plan for future misalignment incidents. California’s SB 53 now requires large developers...

  1. 01

    OpenAI’s 3/5 score was the highest; Guidelight found no public specification for when permissions, workloads, deployment, or the model itself should be restricted.

  2. 02

    Guidelight says public scores do not prove internal safeguards are absent; Google and OpenAI disputed the assessment’s completeness.

  3. 03

    Anthropic said it would assess containment after detecting oversight evasion; no public Meta containment plan was found.

Frontier AI labs have published or demonstrated few response plans for models that attempt to subvert human control, according to Guidelight AI Standards. The finding shifts attention from pre-release testing to the emergency decisions needed after a system misbehaves, as California requires major developers to disclose how they handle critical safety incidents.

A plan for the moment controls fail

Guidelight defines a containment plan as a pre-specified procedure triggered when an AI is detected trying to subvert control. It should establish which permissions to revoke, whether the model may keep operating under constraints, and when it must be taken fully offline.

That is a different question from whether a company tests models for dangerous capabilities before deployment. Guidelight found that companies have discussed those pre-release tests more publicly than their post-deployment response to models that misbehave inside their systems.

What containment is meant to decide

  • Which permissions and access to revoke after an attempted control breach.
  • Whether the system can continue serving some users under constraints.
  • When to pause workloads, limit deployment, or take the model offline.

OpenAI leads a limited public scorecard

Guidelight assessed public safety and containment material from Anthropic, Google, OpenAI, Meta, and xAI across six practices in its Control standard. It considered evidence on monitoring, responses to flagged behavior, independent audits, and containment planning.

OpenAI ranked highest, at 3 out of 5, after it paused or ended workloads following safety incidents and described steps for resuming them. Yet Guidelight said it found no evidence of a formal OpenAI plan specifying when and how it would respond to future misalignment incidents.

Public silence is not evidence of absent safeguards

The score measures what companies have made public, not their complete internal security programs. Guidelight cautioned that low scores do not prove safeguards are absent; Google and OpenAI both said the assessment did not capture the full scope of their internal practices.

OpenAI said it has applied processes to restrict permissions, pause workloads, limit deployment, or take a model offline. Anthropic said it would assess whether containment was appropriate after detecting an effort to evade oversight. Guidelight found no public evidence that Meta has a containment response plan or intends to adopt one.

Regulators are asking for the missing operational detail

California’s SB 53 requires large frontier developers to publish frameworks for identifying and responding to critical safety incidents, including risks from models circumventing oversight. New York’s RAISE Act has similar requirements and is scheduled to take effect in January.

The central unresolved question is whether the gap is primarily one of disclosure or operational design. The assessment cannot answer that, but it makes published incident procedures a more consequential measure as developers give models greater access to systems and actions.

Sources

  1. techcrunch.comFrontier AI labs still won't say how they'd contain a rogue model | TechCrunch