Model
Model explorer

GLM-5.3

OPEN
Zhipu AI / Z.ai (Tsinghua) · GLM-5 family · released Aug 14, 2026

Z.ai's August 14, 2026 release, and an unusually clean natural experiment: the same base model as GLM-5.2, with every gain from post-training alone. Terminal-Bench 3.0 goes 4.6 → 28.3, DeepSWE v1.1 46.2 → 66.9, AutomationBench 26.2 → 48.2, and Z.ai's in-house Code Bench improves ~50% — while consuming FEWER output tokens at every effort level (34.5% at ~75K tokens/task at max, against GLM-5.2's 23.4% at 96K). API-only at launch at $1.40/$4.40 per Mtok with off-peak billing at half rate; Z.ai said weights would follow roughly two weeks later once safety hardening completed, without naming a licence in advance — so `openness` reflects the commitment, and the licence field records that it was still unnamed. The headline finding is emergent cyber capability that Z.ai says developed faster than expected: SOTA on CyberGym at 84.5 (ahead of Mythos 5's 83.8 and GPT-5.6 Sol's 83.6) and more than double GLM-5.2 on exploitation, but the gap to the closed frontier WIDENS further up the exploitation chain — ExploitBench 54.4 against Mythos 5's 78.0 and Sol's 76.5. Its Agents' Last Exam figure (28.5) is omitted: Z.ai reports a pass-rate-style metric incompatible with the Score convention used by the other rows in that file.

ReasoningCodingVisionFunction callingTool useAgentic
3029.9
Elo · rank #6
Parameters
744B
Active params
40B (MoE)
Context
200K tokens
Architecture
The SAME base model as GLM-5.2 — every gain comes from scaled post-training (IndexShare long-context, SAO long-horizon RL, slime async training). Sparse MoE, three thinking-effort levels (low/high/max); disabling thinking is no longer supported
License
Weights published on Hugging Face two weeks after the API launch; licence not named in advance by Z.ai
Languages
API price (in/out)
$1.4 / $4.4
Modalities
text
Benchmark results
Bar shows position within the tracked field; marker = field best
AutomationBenchAgents48.2%#4
best: Qwen3.8-Max-0902 · 50.8%
CyberGymCoding84.5%#1
best: this model · 84.5%
DeepSWECoding66.9%#7
best: Muse Spark 1.3 · 75.4%
ExploitBenchAgents54.4%#3
best: GPT-6 Astra · 100.0%
NL2RepoCoding58.0%#1
best: this model · 58.0%
Terminal-Bench 2.0Coding88.2%#5
best: Gemini 3.8 Flash · 89.4%
Terminal-Bench 3.0Agents28.3%#3
best: GPT-5.6 Sol · 34.6%
best: GLM-5.3-Flash · 78.4%
Run it locally
VRAM @ Q4
VRAM @ FP16
Fits on (Q4)
Multi-node cluster required
Throughput data unavailable.
Quantizations
Fine-tune it
Conditional / custom
QLoRA505.9 GB4× H200 141GB
LoRA1584.7 GBbeyond 8× B200
Full fine-tune11941.2 GBbeyond 8× B200
QLoRA SFT on ~10k samples ≈ $20.69 (4× H200 141GB)
API price $1.4/$4.4 · each benchmark row carries its own source badge (see methodology)