Baseten’s Base Labs Partners With Hugging Face and Goodfire on Open-Model Safety

The partners want safety checks and monitoring built into open-model development, but have not explained how the system will work.

By 3 min read
Baseten’s Base Labs Partners With Hugging Face and Goodfire on Open-Model Safety
Baseten’s Base Labs Partners With Hugging Face and Goodfire on Open-Model Safety

Listen to this story

The audio brief

About 1:21
0:001:21
Read transcript
Base Labs, Baseten’s research arm, is partnering with Hugging Face and Goodfire AI on a proposed safety framework for open-weight AI models. The plan is to build safety checks into the entire model lifecycle, from training and evaluation through monitoring and deployment, rather than adding safeguards after a model is already finished. The partnership responds to the spread of “abliterated” models—models whose safety measures have been removed. Hugging Face lists more than six thousand of them, a figure that illustrates the scale of the problem, but does not show whether this new framework could detect or prevent those changes. So far, the announcement describes a direction, not a working system. The companies have not said which evaluations, monitoring tools, or training techniques they will publish. They also have not explained what developers would need to do to comply, or how compliance would be independently assessed. That matters because the effort is being presented as an open standard, not a proprietary safety layer. Base Labs says it will publish its methods and invite contributions from the broader developer community. Baseten brings experience as an AI inference provider, while Goodfire focuses on making model behavior more interpretable. But participation alone will not establish a standard. The key question is whether outsiders can clearly inspect the methods, measure results, and judge whether an open model actually meets the framework’s requirements.

Story brief

3 key points

Base Labs, Baseten’s research arm, is working with Hugging Face and Goodfire AI on a proposed open-model safety framework spanning training, evaluation, monitoring, and deployment. The effort responds to the growing availability of “abliterated” models—Hugging Face lists more than 6,000—but the partners have not explained how their system would detect or prevent safeguard removal. Base Labs plans to publish methods...

  1. 01

    Hugging Face lists more than 6,000 abliterated models, illustrating the safety problem the partnership aims to address.

  2. 02

    Base Labs has not disclosed the evaluations, monitoring systems, or training techniques it will publish.

  3. 03

    The proposed effort is framed as an open standard, not a proprietary safety layer.

Base Labs is promising a transparent safety framework for open-weight AI models that begins during training and continues through deployment. Its new partnership with Hugging Face and Goodfire AI offers an answer to a real concern around removable safeguards, but the companies have not yet disclosed the technical structure behind that promise.

Baseten’s research arm will work with Hugging Face and Goodfire on safety evaluation and monitoring infrastructure for open-weight models. Base Labs also says it will develop and publish methods for training and monitoring open models, positioning the effort as a proposed standard rather than a proprietary safety layer.

Safety before and after release

The stated aim is to make the framework transparent and part of model training and deployment, instead of applying it after a model has been built. That distinction goes to the partnership’s central premise: safety practices should accompany an open model through its lifecycle, not arrive as a later add-on.

The need for durable safeguards is sharpened by abliterated models, a term used for models whose safety measures have been removed. Hugging Face currently lists more than 6,000 such models, according to TechCrunch. That count does not show whether the proposed framework can prevent or detect those changes, but it illustrates the environment the partners are addressing.

Abliterated models on Hugging Face
More than 6,000Abliterated models listed on Hugging Face

TechCrunch reports that Hugging Face currently lists more than 6,000 abliterated models.

A framework without an implementation

The announcement supplies a direction, not yet an operating blueprint. The companies have not disclosed how the partnership will work technically. It therefore remains unclear which evaluations, monitoring practices or training methods the partners will publish, and how those pieces would be used by developers or providers.

That gap matters for the word “standard.” A standard can describe a shared set of methods, a way to judge safety results, or an expectation for how models are served. Base Labs has committed to publishing methods, but it has not said what technical requirements participants would have to meet or how compliance would be assessed.

An invitation beyond the three companies

Base Labs has invited the broader developer ecosystem to contribute to the framework. That makes the effort an attempt to gather participation around published practices, not simply a bilateral integration among infrastructure vendors. Whether that invitation produces a broadly used framework will depend on details the announcement does not yet provide.

For Baseten, the partnership also extends the remit of Base Labs, the research group it launched earlier this year. Baseten is an AI inference provider, while Goodfire focuses on making model behavior more interpretable. The partnership puts those organizations alongside Hugging Face around a shared goal, but its credibility will rest on the methods they eventually release and how clearly outsiders can evaluate them.

Sources

  1. techcrunch.comBase Labs launches an open-weight AI safety partnership with Hugging Face and Goodfire | TechCrunch

Loading discussion...

Baseten’s Base Labs Partners With Hugging Face and Goodfire on Open-Model Safety | Superpower Daily