Mostik Links AI Models Through Their Weights, Claiming a 20x Cost Cut
The startup is proposing an alternative to the usual model-to-model text handoff. Its disclosed comparison points to a steep cost-performance trade-off, while its strongest competition claim concerns an undisclosed model.
Listen to this story
The audio brief
Story brief
3 key pointsMostik’s disclosed experiment suggests a cheaper way to combine models, but not a breakthrough in parity: a 4B Qwen-3.5 model was linked with 753B GLM-5.2 through model weights rather than text. The company claims the hybrid costs one-twentieth as much as GLM-5.2 and delivers midpoint performance, though it did not identify the evaluation. The approach could make specialized and open-weight models more useful, while...
- 01
The disclosed pairing spans 753B and 4B parameters; Mostik says Qwen-3.5 can run on a mobile device.
- 02
Weight-level communication aims to avoid the generation and processing overhead of conventional model-to-model handoffs.
- 03
Midpoint performance indicates an efficiency trade-off, not GLM-5.2 parity; the benchmark remains undisclosed.
Mostik has demonstrated a system that lets AI models exchange information through internal weights instead of generated text. It paired a 753-billion-parameter GLM-5.2 model with a 4-billion-parameter Qwen-3.5 model, claiming one-twentieth the cost of the full GLM model and performance halfway between the two.
Skipping the text relay
Model weights are the mathematical values that determine how a prompt becomes an output. Mostik’s approach uses those values as the connection between models, rather than making one generate text for the next to process.
That differs from a conventional ensemble, where models are combined by feeding one model’s output into another. Such handoffs can improve results, but the generation and processing add time and cost. Mostik CEO Sasha Malysheva developed the technique.
A cost saving, not large-model parity
The disclosed test joins models at very different scales. Mostik used the largest GLM-5.2, with 753 billion parameters, and a 4-billion-parameter Qwen-3.5 version that can run on a mobile device.
Mostik said the hybrid costs one-twentieth as much as the full GLM model, but its performance landed exactly halfway between the larger and smaller models. That makes the example a stated efficiency trade-off, not a claim that the cheaper system equals GLM-5.2. The disclosed comparison does not identify the evaluation behind that performance result.
The case for specialized models
Malysheva argues that future AI capability need not come from a single ever-larger model. The technique could raise the value of open-weight models by helping them compete with closed proprietary systems.
Vladimir Arustamian, tech lead at Lovable, said a bridge between frontier and domain-specific models could lead to more specialized systems being trained. But Mostik chief scientist Stanislav Smirnov says finding common ground between two models is difficult, with no appropriate mathematical language yet.
A withheld benchmark result
Mostik also says an undisclosed model built with its method reached the top of the ARC-AGI 3 competition. The company declined to provide more information because it wants to win the contest, leaving the GLM-Qwen hybrid as the clearest disclosed example of the bridge’s claimed cost and capability balance.
Sources
- wired.comThese Russian Mathematicians Taught AI Models How to Talk to Each Other Without Using Words