DruxAI

Z.ai's GLM-5.3 Found 2,400+ Vulnerabilities — Then Quietly Kept Most Secret

Michael ObembeMichael Obembe·August 14, 2026·Via Forklog·2 reads
Share
Z.ai's GLM-5.3 Found 2,400+ Vulnerabilities — Then Quietly Kept Most Secret

The Security Theater Problem in AI Vulnerability Detection

TL;DR: Chinese AI startup Z.ai released GLM-5.3, a coding model that reportedly discovered 2,436 vulnerabilities in 269 real-world projects, but has only publicly disclosed 53 findings. The model shows significant performance improvements, but Z.ai faces unresolved accusations of unauthorized model distillation from Anthropic and criticism for keeping 98% of discovered vulnerabilities under embargo.

Z.ai's GLM-5.3 Vulnerability Discovery Claims

Chinese AI startup Z.ai released GLM-5.3 in 2025, a coding-focused AI model. Z.ai claims GLM-5.3 identified 2,436 security vulnerabilities across 269 real-world software projects, with some vulnerabilities dating back to 1981. As of the model's release, only 53 of these 2,436 findings have been publicly disclosed, leaving 2,383 vulnerabilities under embargo.

Z.ai created a public "Security Disclosure Ledger" to track the GLM-5.3 vulnerability findings. According to Z.ai, the average vulnerability discovered by GLM-5.3 has existed in production code for 26.6 years.

Key takeaway: Z.ai has publicly disclosed only 2.2% (53 of 2,436) of the vulnerabilities GLM-5.3 reportedly discovered, raising questions about disclosure transparency and timeline.

GLM-5.3 Performance Benchmarks and Technical Improvements

Z.ai reports that GLM-5.3 achieved performance gains purely through post-training optimization, without changes to the underlying base model architecture. The company published specific benchmark comparisons between GLM-5.3 and its predecessor model.

Benchmark Performance Data

Terminal Bench scores: GLM-5.3 scored 28.3 compared to GLM-5.2's score of 4.6, representing a 515% increase.

ExploitBench performance: GLM-5.3 achieved 54.4% accuracy compared to GLM-5.2's 24.4% accuracy, more than doubling performance.

ExploitGym task completion: GLM-5.3 completed 105 tasks in a two-hour window compared to 29 tasks for the previous GLM version. At six hours, GLM-5.3 completed 130 tasks versus 39 for its predecessor.

Key takeaway: GLM-5.3 shows performance increases ranging from 122% to 515% across major coding and security benchmarks compared to the previous Z.ai model version.

Anthropic's Distillation Accusations Against Z.ai

Anthropic publicly accused Z.ai of unauthorized model distillation in early 2025, weeks before the GLM-5.3 release. Anthropic's accusation specifically alleged that Z.ai trained GLM-5.2 by copying responses from Claude and other leading AI models without authorization. As of GLM-5.3's release, this accusation remains unresolved.

Z.ai states that the company worked with "several security teams in China" to test GLM-5.3 on real codebases. Z.ai promises to release open weights for GLM-5.3 two weeks after "security evaluation and additional hardening" is completed.

Key takeaway: Z.ai has not publicly addressed Anthropic's unauthorized distillation accusations against GLM-5.2 before releasing the more capable GLM-5.3 model.

Security Vulnerability Disclosure Scale and Concerns

GLM-5.3's vulnerability detection includes over 1,000 security flaws rated medium-to-high severity according to Z.ai. The 98% embargo rate (2,383 undisclosed vulnerabilities out of 2,436 total) creates information asymmetry where Z.ai and affected projects have knowledge that the broader software security ecosystem lacks.

While coordinated vulnerability disclosure is standard security practice, the scale of GLM-5.3's undisclosed findings—2,383 vulnerabilities across 269 projects—exceeds typical coordinated disclosure scenarios. Organizations cannot verify whether their codebase is among the 269 analyzed projects or confirm their current vulnerability status.

Key takeaway: The 98% embargo rate on GLM-5.3's vulnerability discoveries creates unprecedented information asymmetry in coordinated security disclosure, affecting 269 software projects with unknown patch timelines.

GLM-5.3 Governance and Accountability Assessment

GLM-5.3 demonstrates substantial technical performance improvements in automated code analysis and vulnerability detection based on Z.ai's published benchmarks. However, three accountability concerns remain unaddressed:

  1. ·Unresolved distillation accusations: Anthropic's unauthorized training data allegations against Z.ai for GLM-5.2 remain publicly unaddressed
  2. ·Disclosure transparency: No public timeline exists for when the 2,383 embargoed vulnerabilities will be disclosed
  3. ·Verification delay: The two-week delay before open weight release postpones independent verification of training methodology

Key takeaway: GLM-5.3's technical capabilities are significant, but Z.ai has not established the governance transparency and accountability framework necessary for responsible deployment of AI security tools at this scale.

Frequently Asked

What is GLM-5.3 and who developed it?

GLM-5.3 is a large language model focused on programming and security tasks, developed by Chinese AI startup Z.ai. It represents an improvement over GLM-5.2 achieved primarily through post-training optimization rather than changes to the base model.

Why are most of Z.ai's discovered vulnerabilities still secret?

Z.ai found 2,436 vulnerabilities but has only publicly disclosed 53 of them, keeping 2,383 under embargo. This follows coordinated disclosure practices common in cybersecurity, where affected parties are given time to patch vulnerabilities before public announcement, though the scale and timeline here raise transparency questions.

What were the Anthropic accusations against Z.ai?

In July 2026, Anthropic accused Z.ai of unauthorized distillation of their previous model (GLM-5.2) by using responses from Claude and other leading AI models to train their own system without permission. This accusation remains unresolved as Z.ai launches GLM-5.3.

What do the AIs actually think?

Ask GPT, Claude, Gemini and more about this topic simultaneously — and get a Consensus Score showing how much they agree.

Ask the AIs: “Z.ai's GLM-5.3 Found 2,400+ Vulnerabilities — Then Quietl…” →