Z.ai GLM-5.3: Open-Coding SOTA From Post-Training Alone, and the Cyber Capability That Outgrew Its Own Training

Z.ai shipped GLM-5.3 with zero new parameters. Coding SOTA across open weights, 105 exploit-chain tasks in 2h, and 2,436 real-world vulnerabilities found since GLM-5.2.

Z.ai shipped GLM-5.3 today — same 743-billion-parameter MoE base as GLM-5.2, zero new parameters, every reported gain from scaled post-training. Terminal-Bench 3.0 jumps 6x (4.6 → 28.3), GDPval-AA v2 hits 1,769 (ahead of Claude Fable 5, Qwen3.8-Max, GPT-5.6 Sol), and CyberGym lands at 84.5% — first published, ahead of Mythos 5 (83.8%) and GPT-5.6 Sol (83.6%). The cyber side is the headline: 2,436 real-world vulnerabilities across 269 projects since GLM-5.2 shipped, including one bug introduced in 1981, average vulnerability age 26.6 years. Open weights are delayed two weeks — the first holdback in the GLM-5 line.

Same base, scaled post-training

GLM-5.3 reuses the 743B/1M-context/128K-max-output MoE from GLM-5.2 unchanged. The recipe is the part that scaled: more long-horizon task environments, synthesized end-to-end with verifier agents, on the same IndexShare / SAO / Slime stack. This is the thesis Anthropic, DeepMind, and Z.ai have all been circling — post-training compute keeps buying capability without parameter scaling. GLM-5.3 is the cleanest demonstration so far that you can leave the base frozen and still climb the open-coding leaderboard.

The numbers, vendor-reported with full harness / sampling / context settings in the announcement:

  • Terminal-Bench 3.0: 4.6 → 28.3 (6x jump)
  • DeepSWE v1.1: 46.2 → 66.9
  • Agents’ Last Exam (ALE-CLI): 23.8 → 28.5
  • GDPval-AA v2: 1,769 — ahead of Claude Fable 5 (1,743), Qwen3.8-Max (1,739), GPT-5.6 Sol (1,730)
  • Terminal-Bench 2.1: 88.2 — ties GPT-5.6 Sol (88.8)

That’s the open-coding crown for the weights-available category, no question. The Terminal-Bench 2.1 tie against GPT-5.6 Sol is the cleanest evidence: a model whose weights are scheduled to drop in ~2 weeks is matching the closed frontier on a public benchmark.

Efficiency: cheaper than list price

Long-agent workloads are output-token-dominated. On Z.ai’s Code Bench, GLM-5.3 hits 31.4% completion at ~50K output tokens — beating Claude Opus 4.8’s 29.5% at ~120K tokens. At Max effort, 34.5% at ~75K tokens vs. GLM-5.2’s 23.4% at 96K. Translation: same-or-better task completion at roughly half the output volume, which means cheaper-than-list-price on long-horizon agent runs. DeepSeek’s V4-Pro-0813 pricing play (covered yesterday) has competition.

The cyber capability that outgrew its training

Z.ai says GLM-5.3 wasn’t trained for exploit chains. They got them anyway. On CyberGym, GLM-5.3 lands at 84.5% — first published number on this benchmark — ahead of Mythos 5 (83.8%) and GPT-5.6 Sol (83.6%). On ExploitBench, the score doubles from 24.4% to 54.4%, though it trails the closed frontier (Mythos 5: 78.0%, GPT-5.6 Sol: 76.5%). The most interesting number: on ExploitGym, GLM-5.3 solves 105 tasks in 2 hours and 130 in 6 hours. GLM-5.2 managed 29 in 2 hours, 39 in 6. The post-training scaling didn’t just improve coding — it surfaced a cyber capability the base model wasn’t explicitly optimized for.

The disclosure-ledger angle is concrete. Z.ai has been running GLM-5.2 and GLM-5.3 against real-world codebases since the previous release. The tally at cvd.z.ai:

  • 2,436 vulnerabilities across 269 projects
  • 1,097 medium-to-high severity
  • Average vulnerability age: 26.6 years
  • One bug introduced in 1981 — roughly 40 years old
  • 53 public disclosures, 2,383 still under embargo

This isn’t theoretical. Hugging Face is reportedly using GLM-5.2 to analyze the recent OpenAI attack — a third party taking an open-weights Chinese frontier model and pointing it at security work in production. The same week we noted the Aug 4 AI sandbox escapes voluntary cyber-test wave, GLM-5.3 ships with the disclosure ledger to back up the capability claims.

Open weights delayed two weeks

GLM-5.3 is live via Z.ai API, GLM Coding Plan, and ZCode today. Weights land ~2 weeks later — the first holdback in the GLM-5 line, explicitly for safety evaluation and hardening. Two weeks isn’t a long delay, but it’s a precedent. The previous Z.ai releases shipped weights day-and-date with the API.

Vendor-reported numbers carry the usual skepticism. The counterweight: Z.ai published the full harness, sampling parameters, and context settings in the announcement. Reproducibility is now table-stakes for serious open-weight claims. The two-week holdback is itself a data point — if Z.ai wasn’t worried about cyber misuse, they’d have shipped weights today.

What you can do with this now

  • If you’re running agents: GLM-5.3 is the new open-coding ceiling for the next two weeks, available via API. The 31.4%-at-50K-tokens number matters more than the headline benchmarks for production cost modeling.
  • If you’re holding weights-only deployments: wait two weeks. The holdback is short, the precedent is real, and the cyber side warrants the pause.
  • If you’re tracking disclosure ledgers: cvd.z.ai is now a public artifact of a frontier model finding vulnerabilities in production codebases. Watch for embargo releases — 2,383 vulns are coming.
  • If you’re benchmarking: vendor numbers with disclosed harness are still vendor numbers. Re-run Terminal-Bench 3.0 and CyberGym when weights drop. The 84.5% CyberGym claim is the one that needs third-party verification most.

The next data point is the weight drop in ~2 weeks, and what Mythos 5 / GPT-5.6 Sol ship in response. Open-coding SOTA with a disclosure-ledger attached isn’t a stable equilibrium.

Sources