Modelspublished

Mostik Links AI Models Through Their Weights, Claiming a 20x Cost Cut

The startup is proposing an alternative to the usual model-to-model text handoff. Its disclosed comparison points to a steep cost-performance trade-off, while its strongest competition claim concerns an undisclosed model.

By 2 min read
Mostik Links AI Models Through Their Weights, Claiming a 20x Cost Cut
Mostik Links AI Models Through Their Weights, Claiming a 20x Cost Cut

Listen to this story

The audio brief

About 1:25
0:001:25
Read transcript
Mostik says it has linked two AI models through their internal weights, avoiding the usual text handoff between models. In its clearest disclosed test, a 753-billion-parameter GLM-5.2 model was connected to a much smaller, 4-billion-parameter Qwen-3.5 model—small enough, Mostik says, to run on a mobile device. The company says this hybrid costs one-twentieth as much as the full GLM-5.2, while delivering performance halfway between the larger and smaller models. That is a significant efficiency trade-off, but not parity with the largest model; Mostik has not disclosed which evaluation produced the result. The technical idea is that model weights—the mathematical values that shape how a model turns a prompt into an answer—can serve as the connection. That could remove the generation and processing overhead of having one model write text for another, as in a conventional ensemble. The broader pitch is that specialized and open-weight models could become more useful alongside frontier systems, instead of capability depending entirely on one enormous model. But Mostik chief scientist Stanislav Smirnov says finding a shared mathematical language between models remains difficult. The company also claims an undisclosed system reached the top of the ARC-AGI 3 competition, while withholding details. So the key constraint is still proof: can this weight-level bridge scale beyond one undisclosed benchmark and one carefully chosen pairing?

Story brief

3 key points

Mostik’s disclosed experiment suggests a cheaper way to combine models, but not a breakthrough in parity: a 4B Qwen-3.5 model was linked with 753B GLM-5.2 through model weights rather than text. The company claims the hybrid costs one-twentieth as much as GLM-5.2 and delivers midpoint performance, though it did not identify the evaluation. The approach could make specialized and open-weight models more useful, while...

  1. 01

    The disclosed pairing spans 753B and 4B parameters; Mostik says Qwen-3.5 can run on a mobile device.

  2. 02

    Weight-level communication aims to avoid the generation and processing overhead of conventional model-to-model handoffs.

  3. 03

    Midpoint performance indicates an efficiency trade-off, not GLM-5.2 parity; the benchmark remains undisclosed.

Mostik has demonstrated a system that lets AI models exchange information through internal weights instead of generated text. It paired a 753-billion-parameter GLM-5.2 model with a 4-billion-parameter Qwen-3.5 model, claiming one-twentieth the cost of the full GLM model and performance halfway between the two.

Skipping the text relay

Model weights are the mathematical values that determine how a prompt becomes an output. Mostik’s approach uses those values as the connection between models, rather than making one generate text for the next to process.

That differs from a conventional ensemble, where models are combined by feeding one model’s output into another. Such handoffs can improve results, but the generation and processing add time and cost. Mostik CEO Sasha Malysheva developed the technique.

A cost saving, not large-model parity

The disclosed test joins models at very different scales. Mostik used the largest GLM-5.2, with 753 billion parameters, and a 4-billion-parameter Qwen-3.5 version that can run on a mobile device.

Mostik said the hybrid costs one-twentieth as much as the full GLM model, but its performance landed exactly halfway between the larger and smaller models. That makes the example a stated efficiency trade-off, not a claim that the cheaper system equals GLM-5.2. The disclosed comparison does not identify the evaluation behind that performance result.

The case for specialized models

Malysheva argues that future AI capability need not come from a single ever-larger model. The technique could raise the value of open-weight models by helping them compete with closed proprietary systems.

Vladimir Arustamian, tech lead at Lovable, said a bridge between frontier and domain-specific models could lead to more specialized systems being trained. But Mostik chief scientist Stanislav Smirnov says finding common ground between two models is difficult, with no appropriate mathematical language yet.

A withheld benchmark result

Mostik also says an undisclosed model built with its method reached the top of the ARC-AGI 3 competition. The company declined to provide more information because it wants to win the contest, leaving the GLM-Qwen hybrid as the clearest disclosed example of the bridge’s claimed cost and capability balance.

Sources

  1. wired.comThese Russian Mathematicians Taught AI Models How to Talk to Each Other Without Using Words