Model
Model explorer

GLM-5

OPEN
Zhipu AI / Z.ai (Tsinghua) · GLM-5 family · released Feb 11, 2026

Major-version flagship succeeding GLM-4.6; first Z.ai model released after Zhipu's Jan 2026 Hong Kong IPO; widely reported as trained entirely on domestic Huawei chips, no Nvidia used.

ReasoningCodingVisionFunction callingTool useAgentic
2096.8
Elo · rank #69
Parameters
744B
Active params
40B (MoE)
Context
200K tokens
Architecture
Mixture-of-Experts, 744B total / 40B active, DeepSeek Sparse Attention (DSA), 200K context, trained on a 100k-GPU Huawei Ascend 910B cluster (no Nvidia/AMD/Intel silicon)
License
MIT
Languages
API price (in/out)
$1 / $3.2
Modalities
text
Benchmark results
Bar shows position within the tracked field; marker = field best
AIMEMath92.7%#34
best: GPT-5.2 · 100.0%
BrowseCompAgents62.0%#29
best: Kimi K3 · 91.2%
CyberGymCoding43.2%#13
best: GLM-5.3 · 84.5%
GPQA DiamondReasoning86.0%#58
best: GPT-6 Astra · 96.0%
Humanity's Last ExamReasoning30.5%#51
best: Claude Opus 5 · 64.7%
IFBenchReasoning72.0%#28
best: MiniMax M3 · 83.0%
LiveCodeBenchCoding81.9%#42
best: DeepSeek-V4-Pro (Think Max) · 93.5%
MCP AtlasAgents67.8%#6
best: Kimi K3 · 84.2%
SWE-bench VerifiedCoding77.8%#29
best: Claude Opus 5 · 96.0%
τ²-Bench TelecomAgents89.7%#21
best: Claude Opus 4.6 · 99.3%
Terminal-Bench 2.0Coding56.2%#52
best: Gemini 3.8 Flash · 89.4%
Run it locally
VRAM @ Q4
VRAM @ FP16
Fits on (Q4)
Multi-node cluster required
Throughput data unavailable.
Quantizations
Fine-tune it
Permissive
QLoRA505.9 GB4× H200 141GB
LoRA1584.7 GBbeyond 8× B200
Full fine-tune11941.2 GBbeyond 8× B200
QLoRA SFT on ~10k samples ≈ $20.69 (4× H200 141GB)
API price $1/$3.2 · each benchmark row carries its own source badge (see methodology)