Z.AI Adds GLM-5.3-Flash to Coding Plan but Leaves FlashX Off It

The release pairs multimodal, long-context capabilities with two distinct access options: a higher-quota plan model and a faster API variant.

By 2 min read
Z.AI Adds GLM-5.3-Flash to Coding Plan but Leaves FlashX Off It
Z.AI Adds GLM-5.3-Flash to Coding Plan but Leaves FlashX Off It

Listen to this story

The audio brief

About 1:30
0:001:30
Read transcript
Z.AI is giving Coding Plan subscribers three times the GLM-5.3 quota with its new GLM-5.3-Flash model—but not with the faster FlashX variant. That access split is the practical headline: both models are available through Z.AI’s API, yet only Flash is included in the subscription plan for now. Flash is positioned as the first native multimodal model in the GLM-5 line. It can take video, images, text, and files, and it supports a context window of one million tokens—enough for very large codebases or document-heavy workflows. Z.AI says the model has 320 billion total parameters, with 18 billion active for any given task. The company is also pitching Flash as more efficient on long inputs. Compared with GLM-5.3, Z.AI reports 3.01 times less attention computation and 4.44 times lower key-value-cache usage. The target use cases range from visual coding and browser interaction to document processing, financial research, video understanding, Blender scenes, games, and CAD generation. FlashX offers a different advantage: Z.AI advertises inference at 200 tokens per second. But that speed comes with a constraint beyond plan access. Thinking mode is enabled by default and cannot be turned off on either variant. So the immediate question is not simply which model is faster. It is whether the Coding Plan’s extra capacity matters more than FlashX’s advertised speed—and how long FlashX remains API-only.

Story brief

3 key points

Z.AI’s latest Flash model changes the value proposition of its Coding Plan: GLM-5.3-Flash is included with three times GLM-5.3’s quota, while the faster FlashX variant remains API-only for now. Flash supports multimodal inputs and a one-million-token context, with Z.AI reporting substantially lower attention and cache costs for long prompts. Developers therefore face a tradeoff between plan-based capacity and...

  1. 01

    GLM-5.3-Flash receives three times GLM-5.3’s Coding Plan quota; FlashX is excluded.

  2. 02

    Both models are available through Z.AI’s API as `glm-5.3-flash` and `glm-5.3-flashx`.

  3. 03

    Flash accepts video, images, text, and files, with a one-million-token context window.

Z.AI’s GLM-5.3-Flash release gives coding-plan subscribers a larger included allowance, but its faster FlashX sibling is not part of that plan. Both models are available through Z.AI’s API, making access terms—not just model speed—the immediate difference between the variants.

Z.AI describes Flash as the first native multimodal model in its GLM-5 series. It accepts video, images, text and files, produces text, and has a one-million-token context window. The model has 320 billion total parameters, with 18 billion activated parameters.

For developers calling the models directly, Z.AI lists the identifiers glm-5.3-flash and glm-5.3-flashx. Vercel AI Gateway also lists FlashX under zai/glm-5.3-flashx, with a one-million-token context window and a maximum output limit of 131,072 tokens.

An efficiency pitch for long inputs

Z.AI says Flash combines sparse and linear attention, a design it positions as a way to reduce the work of handling long context. Compared with GLM-5.3, the company reports a 3.01-fold reduction in attention computation and a 4.44-fold reduction in key-value-cache usage.

The workflows Z.AI is targeting

  • Visual coding and interaction across browsers and graphical interfaces.
  • Office-document generation, document processing and financial research workflows.
  • Video understanding, Blender scenes, game development and CAD-model generation.

The documentation also lists function calling, structured JSON output, streaming responses and context caching. Its thinking mode is enabled-only, so developers cannot turn that mode off. The release’s practical appeal will depend on whether the plan-based Flash option or FlashX’s stated response speed better fits a given workflow.

Sources

  1. vercel.comGLM 5.3 FlashX API, Pricing & Playground | Vercel AI Gateway
  2. docs.z.aiGLM-5.3-Flash/FlashX - Overview - Z.AI DEVELOPER DOCUMENT

Loading discussion...

YOUR READING SPACE

Notifications