Model
Model explorer

GLM-5.1

OPEN
Zhipu AI / Z.ai (Tsinghua) · GLM-5 family · released Apr 7, 2026

Point release focused on extreme long-horizon agentic engineering; Z.ai's own SWE-bench Pro claim (58.4) briefly topped the public leaderboard, edging out GPT-5.4 and Claude Opus 4.6 on that benchmark at launch.

ReasoningCodingVisionFunction callingTool useAgentic
2224.1
Elo · rank #57
Parameters
744B
Active params
40B (MoE)
Context
200K tokens
Architecture
Mixture-of-Experts, 744B total / 40B active, DeepSeek Sparse Attention (DSA), 200K context, tuned for autonomous long-horizon agent runs (up to 8 hours single-run)
License
MIT
Languages
API price (in/out)
$1.4 / $4.4
Modalities
text
Benchmark results
Bar shows position within the tracked field; marker = field best
Agents' Last ExamAgents28.1%#18
best: GPT-6 Astra · 59.3%
AIMEMath95.3%#13
best: GPT-5.2 · 100.0%
BrowseCompAgents68.0%#26
best: Kimi K3 · 91.2%
CyberGymCoding68.7%#8
best: GLM-5.3 · 84.5%
GDPval-AAAgents1260#34
best: Claude Fable 5 · 1932
GPQA DiamondReasoning86.2%#57
best: GPT-6 Astra · 96.0%
Humanity's Last ExamReasoning31.0%#49
best: Claude Opus 5 · 64.7%
IFBenchReasoning76.0%#17
best: MiniMax M3 · 83.0%
LiveCodeBenchCoding81.4%#46
best: DeepSeek-V4-Pro (Think Max) · 93.5%
SWE-bench ProCoding58.4%#22
best: Claude Fable 5.1 · 81.2%
SWE-bench VerifiedCoding76.4%#38
best: Claude Opus 5 · 96.0%
Terminal-Bench 2.0Coding63.5%#39
best: Gemini 3.8 Flash · 89.4%
Run it locally
VRAM @ Q4
VRAM @ FP16
Fits on (Q4)
Multi-node cluster required
Throughput data unavailable.
Quantizations
Fine-tune it
Permissive
QLoRA505.9 GB4× H200 141GB
LoRA1584.7 GBbeyond 8× B200
Full fine-tune11941.2 GBbeyond 8× B200
QLoRA SFT on ~10k samples ≈ $20.69 (4× H200 141GB)
API price $1.4/$4.4 · each benchmark row carries its own source badge (see methodology)