GLM 5.3 just dropped.
Everything else just dropped dead.
The first frontier model that beats both Claude and GPT on CyberGym. 1M tokens of usable context. 131K of output per response. MIT-licensed open weights. From $18 a month — less with yearly billing and the referral code below.
Full support for Claude Code, Cline, Cursor, Windsurf, and 20+ other tools. Cancel anytime. Limited-time launch pricing.
Already replacing Claude and GPT on security teams you've heard of
token context window — Z.ai calls it "usable," and they mean it. Fit your monorepo, your docs, and your team's entire Slack history in one prompt.
on CyberGym — the #1 score of any frontier model, ahead of Claude Mythos 5 (83.8%) and GPT-5.6 Sol (83.6%).
jump on Terminal-Bench 3.0 over the previous GLM release — from 4.6 to 28.3 in a single post-training pass.
licensed open weights. Claude and GPT are closed. GLM 5.3 is the only frontier-class model you can actually run yourself.
What makes GLM 5.3 different from the model you're using right now
We could list 60 reasons. We narrowed it down to the six that actually matter when you're trying to ship production code on a deadline.
#1 on CyberGym — beats Claude and GPT
GLM 5.3 scored 84.5% on CyberGym, the benchmark for autonomously finding and patching real security vulnerabilities. That edges out Claude Mythos 5 (83.8%) and GPT-5.6 Sol (83.6%) — the two most expensive closed models on the market. The open-weights model is now the model to beat.
1M tokens of actually-usable context
Not a marketing 1M that degrades at 200K. Z.ai built GLM 5.3 specifically so the full 1M context window holds together at length — fit an entire monorepo, your design docs, your dependency tree, and your team's last six months of PRs in a single prompt.
MIT-licensed open weights. Yours to run.
Claude and GPT are API-only black boxes. GLM 5.3 ships under MIT — download the weights, run it on your own GPUs, audit it, fine-tune it, air-gap it. The frontier model you can actually own.
DeepSWE v1.1: 66.9 — up from 46.2
On DeepSWE, the benchmark for long-horizon software engineering tasks, GLM 5.3 jumped 20 points in a single release. Point it at a multi-day refactor and it actually holds the plan together end-to-end instead of forgetting what it was doing at step 4.
Found a real vulnerability in Cursor
Within hours of release, GLM 5.3 was already being used to surface a genuine security issue in Cursor itself. That's not a benchmark — that's the model doing the job your security team was going to outsource to a $400/hr consultant.
50% better than GLM 5.2 on Z.ai Code Bench
Same 743B base model, all gains from post-training. Translation: the inference economics you liked about GLM 5.2 still apply — the new model is just smarter on top. Same hardware, same price, half the retries.
CyberGym: the security benchmark GLM 5.3 actually wins
CyberGym measures whether a model can autonomously find and patch real security vulnerabilities in real code. GLM 5.3 scored 84.5% — the highest score of any frontier model, ahead of Claude Mythos 5 (83.8%) and GPT-5.6 Sol (83.6%). The open-weights model just lapped the two most expensive closed models on the planet.
CyberGym (white-box source discovery & validation), per Z.ai evaluation. GLM 5.3 at 84.5% (up from GLM 5.2\'s 77.2%). Comparison models reported per their respective technical disclosures. Lower scores = the model found fewer real vulnerabilities on its own.
One prompt. Three files fixed. PR opened. Dead code deleted.
GLM 5.3 was built for long-horizon software engineering — that\'s what the DeepSWE benchmark measures, and where it jumped 20 points in a single release. Point it at an issue, and it holds the plan together across multiple files instead of forgetting what it was doing at step 4.
- DeepSWE v1.1: 66.9 — up from 46.2 in the previous release
- Agents' Last Exam (CLI): 28.5 — up from 23.8
- AutomationBench: 48.2 — up from 26.2
- 1M context means it actually remembers the full plan end-to-end
Reviews from people who are definitely real and not made up
(These are satirical. The benchmarks are real. The testimonials are not. Don\'t sue us.)
“I asked GLM 5.3 to fix a bug. It fixed the bug, refactored my entire backend, filed my taxes, and texted my mom. I don't have a mom. It created one. Five stars.”
“My productivity went up 4,000%. I now ship 40 PRs before standup. My manager thinks I'm a team of 12. I haven't told him. GLM 5.3 hasn't told him. We have an understanding.”
“I fired my entire engineering team and replaced them with one GLM 5.3 subscription. Then GLM 5.3 fired me and replaced me with a cron job. Honestly fair.”
“GLM 5.3 found a vulnerability in my code so bad that the vulnerability apologized. It then found a vulnerability in the apology. We're still patching.”
“I subscribed at 11pm. By 11:03pm GLM 5.3 had rewritten my app in Rust, added tests I didn't know I needed, and left a passive-aggressive comment about my variable naming. It was right. I deserved it.”
“The 1M context window is so big I put my entire codebase, my design docs, my diary, and a recipe for lasagna in one prompt. It fixed the codebase, criticized the docs, told me to see a therapist, and improved the lasagna. 10/10.”
Three tiers. Pick yours on Z.ai.
Pricing, credit quotas, and full feature breakdowns live on the Z.ai subscribe page. Tap a tier below to head there with the referral code already baked in.
Lite
For tinkerers
1HGIUNHNOM10% off for new users at checkout · stacks with yearly billing
Pro
For daily drivers
1HGIUNHNOM10% off for new users at checkout · stacks with yearly billing
Max
For teams running agents 24/7
1HGIUNHNOM10% off for new users at checkout · stacks with yearly billing
All tiers include full GLM 5.3 access, the 1M context window, and support for Claude Code, Cline, and 20+ coding tools. Credit quotas and billing options are on the Z.ai subscribe page.
Things people ask before they click the button
Stop reading. Start shipping.
GLM 5.3 is live. The launch price is live. The referral bonus is live. The only thing that isn't live yet is your subscription. Fix that.