Thinking Machines Releases Inkling, but Argues Open Weights Need a Gate

The lab’s argument is not that every capable model should be closed. It is that release decisions should turn on whether a model adds risk beyond what is already downloadable—and whether defenders have had time to prepare.

By 3 min read
Thinking Machines Releases Inkling, but Argues Open Weights Need a Gate
Thinking Machines Releases Inkling, but Argues Open Weights Need a Gate

Listen to this story

The audio brief

About 1:38
0:001:38
Read transcript
Thinking Machines Lab has released two downloadable models, Inkling and Inkling-Small, while arguing that open weights should be the final step in a safety process—not the default first move. The lab’s proposed path starts with monitored inference through an API, then hosted fine-tuning, monitored general availability, and only eventually public weights. The logic is straightforward: wider access can help users and defenders work with capable systems, but once weights are published, safety training can be stripped away and access cannot be revoked. For Inkling, the lab applied a narrower, comparative test: would release add material risk beyond what is already available in open-weight models? Internal evaluations covered CBRN, offensive cybersecurity, agentic and tool-use misuse, loss-of-control behavior, and harmful multimodal scenarios. Testing spanned 17 languages, with text, image, and audio inputs. Scale AI, Handshake AI, FAR.AI, and Apollo Research handled external assessments in different risk areas. Thinking Machines also adversarially fine-tuned variants to remove refusal behavior. According to the lab, that produced no new uplift on CBRN or cyber evaluations, leaving results comparable to existing open models. But this is not a universal clearance. The framework is still high-level, especially for more capable systems. Its unresolved issue is operational: what evidence should move a model to the next access stage, and how should defensive readiness be measured? Until those thresholds are published, the proposal is a direction for governance—not a standard others can independently apply.

Story brief

3 key points

Thinking Machines Lab has made Inkling and Inkling-Small downloadable while using the release to propose a conditional alternative to immediate public weights. The lab’s framework would move models through monitored APIs, hosted fine-tuning, monitored general availability and only then open weights, with each step tied to evidence and defensive readiness. Its Inkling assessment—covering 17 languages, multiple...

  1. 01

    Inkling’s evaluation covered CBRN, offensive cyber, agentic misuse, loss of control and harmful multimodal scenarios.

  2. 02

    Scale AI, Handshake AI, FAR.AI and Apollo Research conducted external testing across distinct risk areas.

  3. 03

    Adversarial fine-tuning removed refusal behavior but produced no new CBRN or cyber capability uplift, according to Thinking Machines.

Thinking Machines Lab has released Inkling and Inkling-Small as open-weight language models, while arguing that public weights should be the last step of a safety process rather than the default starting point. Its proposed middle path is to widen access only as evidence supports it and as the surrounding defensive ecosystem becomes more capable.

Publication versus progression

The contrast is between two release logics. Publishing weights lets users own, run and customize a model, but it also makes access irreversible and enables modifications that can remove safety training. Thinking Machines says open models can distribute development and safety work beyond a small group of labs, yet the same distribution creates real misuse risk.

Its alternative is not a fixed ladder that automatically ends in public weights. The lab proposes monitored inference API access, hosted fine-tuning, monitored general availability and, eventually, open weights. Each expansion is meant to give defenders access to capable systems while preserving a provider’s ability to monitor use, maintain guardrails and revoke access before weights are released.

Why Inkling cleared a different bar

For Inkling, the lab did not claim the models were harmless. It framed the decision more narrowly: whether their release would add material incremental risk beyond existing open-weight models. That is a comparative standard, shaped by the lab’s assessment that models of comparable or greater strength in relevant dangerous domains are already downloadable.

The release assessment had three layers

  • Internal tests covered CBRN topics, offensive cybersecurity, broad misuse in agentic and tool-use settings, and multimodal harmful-content scenarios across 17 languages and text, image and audio inputs.
  • Four outside organizations tested distinct areas: Scale AI for general misuse, Handshake AI for vulnerable-user interactions, FAR.AI for CBRN and cybersecurity, and Apollo Research for loss-of-control behaviors.
  • The lab also fine-tuned variants to comply with harmful requests, testing the capabilities exposed after refusal behavior was removed.

Thinking Machines says those tests did not find capabilities that would materially raise real-world risk beyond the existing open-weight baseline. Its helpful-only, adversarially fine-tuned variants produced no new uplift on CBRN or cyber evaluations and remained comparable to existing open-weight models, according to the lab.

A framework designed for a harder future case

The Inkling conclusion does not settle how the framework would work for models nearer the capability frontier. The lab says the balance will change as its systems become more capable, more accessible or easier to modify. Its proposed response is to test dangerous capabilities directly and to build layered defenses before opening access further.

There is a second, more ambitious part of the argument: dangerous specialized knowledge may be separable from general intelligence. Thinking Machines points to pretraining-data curation and post-training interventions as possible ways to reduce dangerous capabilities without broadly degrading useful ones. It treats that possibility as an active research question, not a proven safeguard; a sufficiently capable model might re-derive filtered knowledge, and the boundary between dangerous and ordinary technical knowledge may be unclear.

The decisive unanswered question is operational: what evidence should move a model from one access stage to the next, and how should ecosystem readiness be measured? Thinking Machines calls its post a high-level framework and says it plans a more detailed version covering evaluations and access criteria. Until those thresholds are specified, the proposal is a direction for release governance rather than a release standard others can independently apply.

Sources

  1. thinkingmachines.aiA Safe Path to Open Weights

Loading discussion...