Z.ai Says 100,000 Chinese Chips Serve GLM-5.3-Flash; Shares Rise More Than 8%
The low-cost model is a test of whether China-made hardware can support a public AI service at scale, but Z.ai has not named the chipmakers behind the system.
Listen to this story
The audio brief
Story brief
3 key pointsZ.ai’s GLM-5.3-Flash launch ties a meaningful market reaction to an unverified infrastructure claim: the company says 100,000 China-made chips served every online request, including traffic since the model’s August 20 release. The chip suppliers and exact configuration remain undisclosed. The low-cost model also placed 10th on Artificial Analysis’ Intelligence Index and led OpenRouter usage last week, suggesting...
- 01
GLM-5.3-Flash is a lower-cost version of Z.ai’s flagship model, designed for economical serving rather than training.
- 02
Z.ai says domestic chips handled all online requests, but CNBC could not independently verify the claim or identify suppliers.
- 03
The 100,000-chip figure describes serving capacity for a live service, not the hardware used to train the model.
Z.ai says its new GLM-5.3-Flash model is handling every online request on 100,000 China-made chips, putting a public model service on wholly domestic hardware at substantial claimed scale. Z.ai has not identified the chipmakers, and CNBC could not independently verify the hardware assertions. Its Hong Kong-listed shares rose more than 8% after the launch.
Serving a model is not training one
GLM-5.3-Flash is a low-cost version of Z.ai’s flagship model. Z.ai says the China-made chips handled all online requests, including those from the model’s August 20 release under the code name Ox Alpha.
The claim concerns operating a finished model for users rather than the computing used to train it. Running a model requires less computing power than training one, but a live service still needs capacity ready when users submit prompts. The 100,000-chip figure is therefore a claim about deployment capacity, not simply a demonstration that domestic chips can run the model.
Z.ai says these chips handled all online requests for GLM-5.3-Flash.
Z.ai shares rose more than 8% in Hong Kong trading after the launch.
Model traction does not identify the hardware
The release has model-level traction signals: GLM-5.3-Flash ranked 10th on the Artificial Analysis Intelligence Index, ahead of DeepSeek V4 Pro Max, and ranked first by usage on the global OpenRouter platform during the previous week.
Those measures assess the model’s relative standing and recent use, not the configuration of the hardware serving it. The missing supplier names leave the central infrastructure claim without a public account of which chipmakers are involved.
A domestic alternative under chip restrictions
Nvidia has struggled to sell chips to China because of restrictions from Washington and Beijing, while Huawei and other Chinese companies have increased efforts to build alternatives. China has also expanded domestic semiconductor and AI capabilities in pursuit of technology self-sufficiency after U.S. restrictions on advanced-chip sales.
Leading U.S. AI models are not officially available in China. Against that backdrop, Z.ai’s claim is consequential if substantiated: it describes a domestic hardware base for serving a public AI model, rather than only developing one.
Editorial analysis
Our Read
Our read: The notable part is not merely that Chinese chips can run a model. Z.ai says 100,000 of them are serving the full online workload for a public release. That makes the next useful evidence more operational: identification of the suppliers, or detail that substantiates the serving system’s scale and configuration. The claim arrives while AI companies pursue alternatives to Nvidia-led infrastructure, including OpenAI’s reported custom-chip effort. Z.ai’s release therefore puts attention on whether domestic hardware can support a widely used service, not just a limited demonstration.
Sources
- cnbc.comZ.ai shares surge 8% after releasing new AI model running only on Chinese chips