Modelspublished

Z.ai Says 100,000 Chinese Chips Serve GLM-5.3-Flash; Shares Rise More Than 8%

The low-cost model is a test of whether China-made hardware can support a public AI service at scale, but Z.ai has not named the chipmakers behind the system.

By 3 min read
Z.ai Says 100,000 Chinese Chips Serve GLM-5.3-Flash; Shares Rise More Than 8%

Listen to this story

The audio brief

About 1:27
0:001:27
Read transcript
Z.ai says its new GLM-5.3-Flash model is handling every online request on 100,000 China-made chips—a claim that helped send the company’s Hong Kong-listed shares up more than 8%. The hardware suppliers remain unnamed, and CNBC says it could not independently verify the assertion. The key distinction is serving versus training. GLM-5.3-Flash is a lower-cost version of Z.ai’s flagship model, built for economical serving. Z.ai says domestic chips have handled user traffic since the model’s August 20 release, when it was operating under the code name Ox Alpha. That 100,000-chip figure refers to the hardware supporting a live public service, not the machines used to train the model. Serving generally requires less computing power than training, but it still demands enough capacity to respond as users submit prompts. There are also signs the model is attracting real usage: it ranked 10th on the Artificial Analysis Intelligence Index, ahead of DeepSeek V4 Pro Max, and ranked first by usage on OpenRouter the previous week. The broader backdrop is China’s push for domestic AI hardware as Nvidia faces restrictions on selling chips into the country, while Huawei and other Chinese firms develop alternatives. The unresolved question is whether Z.ai can substantiate the scale—and identify which chipmakers are actually behind it.

Story brief

3 key points

Z.ai’s GLM-5.3-Flash launch ties a meaningful market reaction to an unverified infrastructure claim: the company says 100,000 China-made chips served every online request, including traffic since the model’s August 20 release. The chip suppliers and exact configuration remain undisclosed. The low-cost model also placed 10th on Artificial Analysis’ Intelligence Index and led OpenRouter usage last week, suggesting...

  1. 01

    GLM-5.3-Flash is a lower-cost version of Z.ai’s flagship model, designed for economical serving rather than training.

  2. 02

    Z.ai says domestic chips handled all online requests, but CNBC could not independently verify the claim or identify suppliers.

  3. 03

    The 100,000-chip figure describes serving capacity for a live service, not the hardware used to train the model.

Z.ai says its new GLM-5.3-Flash model is handling every online request on 100,000 China-made chips, putting a public model service on wholly domestic hardware at substantial claimed scale. Z.ai has not identified the chipmakers, and CNBC could not independently verify the hardware assertions. Its Hong Kong-listed shares rose more than 8% after the launch.

Serving a model is not training one

GLM-5.3-Flash is a low-cost version of Z.ai’s flagship model. Z.ai says the China-made chips handled all online requests, including those from the model’s August 20 release under the code name Ox Alpha.

The claim concerns operating a finished model for users rather than the computing used to train it. Running a model requires less computing power than training one, but a live service still needs capacity ready when users submit prompts. The 100,000-chip figure is therefore a claim about deployment capacity, not simply a demonstration that domestic chips can run the model.

The launch in two numbers
100,000Claimed serving hardware

Z.ai says these chips handled all online requests for GLM-5.3-Flash.

More than 8%Hong Kong share move

Z.ai shares rose more than 8% in Hong Kong trading after the launch.

Model traction does not identify the hardware

The release has model-level traction signals: GLM-5.3-Flash ranked 10th on the Artificial Analysis Intelligence Index, ahead of DeepSeek V4 Pro Max, and ranked first by usage on the global OpenRouter platform during the previous week.

Those measures assess the model’s relative standing and recent use, not the configuration of the hardware serving it. The missing supplier names leave the central infrastructure claim without a public account of which chipmakers are involved.

A domestic alternative under chip restrictions

Nvidia has struggled to sell chips to China because of restrictions from Washington and Beijing, while Huawei and other Chinese companies have increased efforts to build alternatives. China has also expanded domestic semiconductor and AI capabilities in pursuit of technology self-sufficiency after U.S. restrictions on advanced-chip sales.

Leading U.S. AI models are not officially available in China. Against that backdrop, Z.ai’s claim is consequential if substantiated: it describes a domestic hardware base for serving a public AI model, rather than only developing one.

Editorial analysis

Our Read

Our read: The notable part is not merely that Chinese chips can run a model. Z.ai says 100,000 of them are serving the full online workload for a public release. That makes the next useful evidence more operational: identification of the suppliers, or detail that substantiates the serving system’s scale and configuration. The claim arrives while AI companies pursue alternatives to Nvidia-led infrastructure, including OpenAI’s reported custom-chip effort. Z.ai’s release therefore puts attention on whether domestic hardware can support a widely used service, not just a limited demonstration.

Sources

  1. cnbc.comZ.ai shares surge 8% after releasing new AI model running only on Chinese chips