Model
Model explorer
GLM-5.2 (Max)
OPENZhipu AI / Z.ai (Tsinghua) · GLM-5 family · released Jun 13, 2026
Highest-capability reasoning-effort tier: raises the thinking-token budget (~85k tokens on hard tasks vs High's ~50% of that) for complex, long-horizon agentic problems. Same weights as the High tier.
ReasoningCodingVisionFunction callingTool useAgentic
2425.7
Elo · rank #40
Parameters
744B
Active params
40B (MoE)
Context
1M tokens
Architecture
Mixture-of-Experts, 744B total / 40B active, DSA + 'IndexShare' sparse-attention indexer reuse (2.9x fewer per-token FLOPs at 1M context), 1M-token context; selectable High/Max reasoning-effort modes
License
MIT
Languages
—
API price (in/out)
$1.4 / $4.4
Modalities
text
Benchmark results
Bar shows position within the tracked field; marker = field best
best: GPT-6 Astra · 59.3%
best: GPT-5.2 · 100.0%
best: GPT-6 Astra · 100.0%
best: Claude Opus 5 · 43.3%
best: Claude Fable 5 · 1932
best: GPT-6 Astra · 96.0%
best: Claude Opus 5 · 64.7%
best: MiniMax M3 · 83.0%
best: DeepSeek-V4-Pro (Think Max) · 93.5%
best: Kimi K3 · 84.2%
best: Claude Fable 5.1 · 81.2%
best: Claude Opus 5 · 96.0%
best: Claude Opus 4.6 · 99.3%
best: Gemini 3.8 Flash · 89.4%
best: GPT-5.6 Sol · 34.6%
best: GLM-5.3-Flash · 78.4%
Run it locally
VRAM @ Q4
—
VRAM @ FP16
—
Fits on (Q4)
Multi-node cluster required
Throughput data unavailable.
Quantizations
—
Fine-tune it
PermissiveQLoRA505.9 GB4× H200 141GB
LoRA1584.7 GBbeyond 8× B200
Full fine-tune11941.2 GBbeyond 8× B200
QLoRA SFT on ~10k samples ≈ $20.69 (4× H200 141GB)
GLM-5 family
Elo progression across releases
API price $1.4/$4.4 · each benchmark row carries its own source badge (see methodology)