Model
Model explorer

GLM-5.3-Flash

OPEN
Zhipu AI / Z.ai (Tsinghua) · GLM-5 family · released Aug 26, 2026

Z.ai's August 26, 2026 Flash-tier model and the first natively multimodal GLM-5. A new base rather than a post-training refresh: 320B total / 18B active with hybrid sparse+linear attention, halving both active parameters and layer count against GLM-4.5's 355B/32B while beating GLM-5.2 across the board at about a tenth of the price. Z.ai places it on the Pareto frontier of Artificial Analysis Intelligence Index v4.1.1 at 57 for $0.045 per task (discounted) — a level previously available only at roughly 10x the cost. Notable for being served entirely on domestic Chinese AI chips, and for having been A/B-tested anonymously as `ox-alpha` on OpenCode and OpenRouter before launch. Pricing here is the published per-Mtok rate; peak/off-peak billing applies on the coding plan. Its HLE row (55.3) is the WITH-tools configuration and its CharXiv row is with tools, both noted per row. Agents' Last Exam is omitted for the same metric-incompatibility reason as GLM-5.3.

ReasoningCodingVisionFunction callingTool useAgentic
3017.6
Elo · rank #7
Parameters
320B
Active params
18B (MoE)
Context
200K tokens
Architecture
Newly trained base: 320B total / 18B active sparse MoE with a hybrid sparse+linear attention stack, 45 layers (against GLM-5.3's 92) — 3.0x less attention compute and a 4.4x smaller KV cache than GLM-5.3. First natively multimodal model in the GLM-5 series
License
Weights published on Hugging Face (zai-org/GLM-5.3-flash); licence not named in the launch post
Languages
API price (in/out)
$0.14 / $0.44
Modalities
text · vision
Benchmark results
Bar shows position within the tracked field; marker = field best
AutomationBenchAgents48.8%#3
best: Qwen3.8-Max-0902 · 50.8%
CharXivVision89.4%#2
best: Qwen3.8-Flash-Next · 90.6%
DeepSWECoding63.4%#10
best: Muse Spark 1.3 · 75.4%
GDPval-AAAgents1773#4
best: Claude Fable 5 · 1932
Humanity's Last ExamReasoning55.3%#5
best: Claude Opus 5 · 64.7%
NL2RepoCoding56.3%#2
best: GLM-5.3 · 58.0%
Terminal-Bench 2.0Coding84.3%#11
best: Gemini 3.8 Flash · 89.4%
best: this model · 78.4%
Run it locally
VRAM @ Q4
VRAM @ FP16
Fits on (Q4)
Multi-node cluster required
Throughput data unavailable.
Quantizations
Fine-tune it
Conditional / custom
QLoRA217.6 GB2× H200 141GB
LoRA681.6 GB4× B200 192GB
Full fine-tune5136.0 GBbeyond 8× B200
QLoRA SFT on ~10k samples ≈ $165 (2× H200 141GB)
API price $0.14/$0.44 · each benchmark row carries its own source badge (see methodology)