Massachusetts Could Make AI Labs Face Public Risk Tests Every 120 Days

The proposal would expose major labs to recurring outside evaluations without giving Massachusetts power to stop development. Its fate now rests with House, Senate and gubernatorial negotiations.

By 3 min read
Massachusetts Could Make AI Labs Face Public Risk Tests Every 120 Days
Massachusetts Could Make AI Labs Face Public Risk Tests Every 120 Days

Listen to this story

The audio brief

About 1:24
0:001:24
Read transcript
Massachusetts lawmakers are considering a rule that would force the biggest AI labs to undergo public, independent risk evaluations at least once every 120 days. The requirement is part of an economic-development bill that cleared the state Senate last month, but it still needs agreement among the House, Senate, and Governor Maura Healey. The rule would apply to labs generating more than 500 million dollars in annual revenue, putting the immediate burden on the industry’s largest developers. Outside evaluators, qualified by the attorney general and paid by the labs, would examine whether frontier models could cause catastrophes: deaths or serious injuries affecting at least 50 people, or property damage of one billion dollars or more. The findings would be public. Labs would also have to disclose safety practices and report incidents. Violations could bring a one-million-dollar initial fine, followed by penalties of up to three million dollars. The debate is partly about independence and partly about speed. Anthropic supports outside evaluations that do not simply check a company’s own safety rules. OpenAI prefers Illinois’s annual third-party audit model, warning that reviews every 120 days could delay cybersecurity-model releases. Massachusetts would gain disclosure and enforcement power, but not the ability to halt development. The key question in the negotiations is whether public reports and fines will provide enough leverage without a release veto.

Story brief

3 key points

A Massachusetts economic-development bill would create a recurring public-review regime for AI labs exceeding $500 million in annual revenue. Independent evaluators—not companies’ own safety teams—would assess risks of deaths, serious injuries, or at least $1 billion in property damage every 120 days. Labs would fund reviews, disclose safety practices and incidents, and face fines up to $3 million, but the state...

  1. 01

    Evaluations would target catastrophic outcomes affecting at least 50 people or causing $1 billion or more in property damage.

  2. 02

    The attorney general would qualify evaluators; labs would provide reasonably necessary materials and pay for assessments.

  3. 03

    Anthropic supports the bill’s independent reviews; OpenAI prefers Illinois’s annual audit model and warns reviews could delay cybersecurity-model releases.

Massachusetts lawmakers are weighing a requirement that major AI labs submit frontier models to independent catastrophic-risk evaluations at least every 120 days, with the findings made public. Anthropic supports the approach; OpenAI says it could delay cybersecurity-model releases and favors a more uniform Illinois-style standard.

The provisions sit inside an economic development bill that cleared the Massachusetts Senate last month. They still need agreement from the House, Senate and Governor Maura Healey. If enacted, the measure would apply only to AI labs with more than $500 million in annual gross revenue, concentrating its immediate obligations on the industry’s largest developers.

A review aimed at catastrophic outcomes

The proposed evaluations go beyond checking whether a company followed its own stated safety process. Outside organizations would assess a model’s potential for catastrophes that kill or seriously injure at least 50 people, or cause $1 billion or more in property damage. The attorney general would set standards for qualifying evaluators, while developers would pay for the reviews.

What the bill would require

  • Outside organizations would conduct catastrophic-risk evaluations at least every 120 days, and their findings would be public.
  • Labs would disclose safety practices and report safety incidents.
  • The state could fine a developer up to $1 million for an initial violation and up to $3 million for later violations involving incident reporting or adherence to its own safety plan.

The dispute is over independence and speed

Anthropic calls the Massachusetts proposal the country’s clearest and strongest AI legislation. Its policy team argues intensive third-party evaluation is necessary because companies should not assess themselves. OpenAI instead backs a uniform standard modeled on Illinois, arguing that inconsistent state rules create confusion rather than added safety.

OpenAI’s more specific objection is operational: it warned that recurring reviews could slow the release of cybersecurity models that might help address the risks lawmakers are trying to manage. The two companies therefore agree on the relevance of external oversight, but differ over its frequency, independence from company rules and possible effect on deployment timelines.

Public visibility without a release veto

The bill’s oversight mechanism is disclosure and compliance enforcement, not direct state control over development. Evaluators could obtain materials reasonably necessary to assess catastrophic risks, but Massachusetts would not gain authority to halt AI development based on their findings. That leaves a central practical question for the final bill: how much leverage public reports and financial penalties will provide when the state cannot block a release.

Negotiators are still working through the AI provisions as part of the larger economic development package. The outcome will determine whether Massachusetts creates a distinct recurring-review regime for large labs or moves closer to the annual audit model OpenAI prefers.

Editorial analysis

Our Read

Massachusetts is testing a sharper version of the state-by-state AI policy debate: whether public disclosure and enforcement can create meaningful oversight without a regulator’s power to stop a model release. The 120-day cadence is the consequential design choice, because it moves evaluation from an annual compliance exercise toward recurring scrutiny of model risk. Anthropic’s support and OpenAI’s resistance show that leading labs do not share a single preferred state framework. Watch whether negotiators preserve the independent-evaluation requirement, and whether the attorney general’s eventual standards define how much access evaluators receive.

Sources

  1. finance.yahoo.comAnthropic, OpenAI Clash Over Strict Massachusetts AI Safeguards
  2. pymnts.comAnthropic and OpenAI Split as Massachusetts Pushes Nation’s Toughest AI Safety Rules | PYMNTS.com

Loading discussion...