Modelspublished

Zhipu Says GLM-5.3 Can Find Bugs Across an Exploitation Chain

The company reports strong vulnerability-discovery results on one benchmark and in 269 real-world projects. It also acknowledges weaker results elsewhere, leaving its wider standing—and the practical force of its access controls—unsettled.

By 4 min read
Zhipu Says GLM-5.3 Can Find Bugs Across an Exploitation Chain

Listen to this story

The audio brief

About 1:43
0:001:43
Read transcript
Zhipu says its new GLM-5.3 can connect separate software weaknesses into a multi-step exploitation chain, rather than just flagging isolated bugs. On the CyberGym vulnerability-discovery benchmark, the company reports that GLM-5.3 outperformed Fable 5 and GPT-5.6 Sol. Zhipu also says it tested the model with Chinese companies across 269 real-world projects, finding 2,436 vulnerabilities. Of those, 1,097 were classified as medium-to-high severity, spanning areas such as operating systems, browser engines, web applications, and network protocols. The important qualification is that this is company-reported evidence. The available accounts do not disclose the model’s false-positive rate, how many findings were fixed, or the precise testing method used against rival systems. Zhipu also acknowledges that GLM-5.3 trails Western models on other security and coding benchmarks. So the release supports a narrower claim: a reported strength in one vulnerability-discovery task, not an independently verified overall lead in cybersecurity or coding. Zhipu is pairing the launch with Shield of Open Source, which offers free security audits and access to automated code-auditing tools through ZCode. At the same time, its Cybersecurity Trusted Access plan is meant to restrict the most sensitive offensive capabilities to verified users, with layered risk reviews for other requests. The open question is whether independent testing confirms both parts of the pitch: the model’s ability to connect exploitation steps, and the practical force of those access controls beyond CyberGym and Zhipu’s own trials.

Story brief

3 key points

Zhipu has launched GLM-5.3 with a cybersecurity pitch centered on linking vulnerabilities into multi-step exploitation chains, rather than detecting isolated bugs. The company reports a CyberGym win over Fable 5 and GPT-5.6 Sol, plus 2,436 vulnerabilities found across 269 real-world projects, including 1,097 medium-to-high-severity issues. However, Zhipu also acknowledges weaker results on other security and coding...

  1. 01

    Zhipu claims GLM-5.3 outperformed Fable 5 and GPT-5.6 Sol on CyberGym vulnerability discovery.

  2. 02

    Testing across 269 projects reportedly produced 2,436 findings, including 1,097 medium-to-high-severity vulnerabilities.

  3. 03

    The evidence is company-reported; false-positive rates, remediation outcomes, and testing methodology remain undisclosed.

Zhipu is presenting GLM-5.3 as a cybersecurity model that does more than spot isolated software flaws. The company says its new system can reason across stages of an exploitation path, a claim that raises the value of automated code review while sharpening the question of who should be allowed to use the same capabilities offensively.

The launch comes with two linked propositions. First, Zhipu says GLM-5.3 outperformed Fable 5 and GPT-5.6 Sol on the CyberGym benchmark for vulnerability discovery. Second, it says the model performed worse than Western models on other security and coding benchmarks. Those statements make the release a narrower claim of strength in one security task, not a demonstrated all-purpose lead in coding or cyber work.

A model claim about the path, not only the flaw

The distinction in Zhipu’s account is between identifying a vulnerability and connecting several steps that could lead from a weakness to its use. The company said GLM-5.3’s cyber capability advanced faster than it expected during post-training, and that the model began reasoning across multiple stages of exploitation. It described the result as coherent plans for complete exploitation chains rather than better detection of isolated flaws alone.

That is the technical mechanism Zhipu is asking customers and developers to evaluate. A system that identifies a single defect may help a team decide what to patch. A system that can join multiple steps in a possible attack sequence could help reviewers understand how separate weaknesses interact. Zhipu’s reported advance is therefore about the connection between findings, though the available reporting does not independently test the quality of those plans.

CyberGym is the centerpiece of the company’s competitive case. Zhipu said GLM-5.3 beat Fable 5 and GPT-5.6 Sol there, but its own acknowledgment of weaker results on other security and coding tests limits what can be inferred from that one result. A benchmark win can establish performance within that test; it does not, on the evidence available here, settle performance across the broader set of security and coding tasks.

Zhipu’s reported real-codebase findings
269Projects tested

Zhipu said it worked with Chinese companies to test GLM-5.3 on real-world codebases spanning 269 projects.

2,436Vulnerabilities found

Zhipu reported that GLM-5.3 found 2,436 vulnerabilities in those real-world codebase tests.

1,097Medium-to-high severity

The company said 1,097 of the reported findings were medium-to-high severity issues.

The field evidence is substantial, but still company-reported

Zhipu also points to tests on code used outside a benchmark. It said it worked with Chinese companies on real-world codebases and found 2,436 vulnerabilities across 269 projects, including 1,097 issues it classified as medium-to-high severity. The reported findings covered system kernels, operating systems, browser engines, open-source infrastructure, web applications and network protocols.

Those figures suggest a broad test set, but they are Zhipu’s figures. The available accounts do not provide an independent assessment of the individual findings, their remediation status, the model’s false-positive rate, or the testing procedure used to compare GLM-5.3 with the named rival models. They also do not reconcile the strong CyberGym result with the company’s statement that the model trails Western systems on other tests.

What Zhipu’s evidence does and does not establish

  • It establishes that Zhipu claims a CyberGym advantage over Fable 5 and GPT-5.6 Sol for vulnerability discovery.
  • It establishes that Zhipu reports real-codebase testing with Chinese companies and thousands of vulnerability findings.
  • It does not establish an independently verified overall lead in cybersecurity or coding, especially because Zhipu says GLM-5.3 performed worse on other benchmarks.

A defensive program beside restricted access

Zhipu launched GLM-5.3 alongside its Shield of Open Source initiative. The program offers free security audits intended to help users patch vulnerabilities, automated code-auditing tools through Zhipu’s ZCode platform, and free model-usage quotas for the open-source community. The structure puts defensive access and developer support beside a more controlled path for the model’s sensitive cyber functions.

For GLM-5.3 itself, Zhipu said it will use a layered risk-review system that blocks high-risk requests while allowing routine, low-risk developer tasks. It also said the most sensitive offensive capabilities will be available only to verified users under a restricted-access plan called Cybersecurity Trusted Access.

The two-part design is consequential because Zhipu’s own framing connects a larger cyber capability with a need for tighter controls. The defensive tools are intended to broaden access to auditing and patching support. The restricted-access plan is intended to narrow access to the most sensitive offensive functions. The reports do not detail how verification works, which capabilities cross that threshold, or how the layered review system will be evaluated after deployment.

The release leaves two tests ahead

One test is technical: whether the model’s claimed ability to connect an exploitation chain holds up beyond CyberGym and Zhipu’s reported company trials. The other is operational: whether Shield of Open Source makes defensive review more available while Cybersecurity Trusted Access reliably constrains the capabilities Zhipu considers most sensitive. The launch supplies the company’s answer to both questions, but not yet independent resolution of either.

Sources

  1. scmp.comZhipu AI’s answer to Project Glasswing marks shift for Chinese cyber safety
  2. theregister.comChinese AI company Zhipu claims its new is a better bug-finder than Anthropic, OpenAI