Model
Model explorer
GLM-4V-9B
OPENZhipu AI · GLM-4 family · released Jun 5, 2024
Open multimodal GLM-4 variant adding 1120x1120 Chinese/English image understanding; claimed to beat GPT-4-turbo-2024-04-09, Gemini 1.0 Pro, Qwen-VL-Max and Claude 3 Opus on multimodal benchmarks.
ReasoningCodingVisionFunction callingTool useAgentic
590.7
Elo · unrated
Parameters
14B
Active params
14B (dense)
Context
8K tokens
Architecture
Dense transformer + ViT vision encoder (GLM-4V, CogVLM-style)
License
GLM-4 License (custom, free for research & most commercial use)
Languages
2+
API price (in/out)
No hosted API
Modalities
text · vision
Benchmark results
Bar shows position within the tracked field; marker = field best
best: Molmo 72B · 96.3%
best: ERNIE 4.5 VL 424B-A47B · 90.9%
best: Qwen3.5-397B-A17B · 93.7%
best: Claude Fable 5 · 89.3%
best: InternVL3-78B · 906
best: NAVER HyperCLOVA X SEED Think 32B · 77.9%
Run it locally
VRAM @ Q4
8 GB
VRAM @ FP16
28 GB
Fits on (Q4)
RTX 3060 12GBRTX 4070 Ti 16GBRTX 3090 24GBRTX 4090 24GBRTX 5090 32GBM4 Pro 48GBM3 Max 128GBM3 Ultra 512GBA100 80GBH100 80GBH200 141GBB200 192GB
Throughput data unavailable.
Quantizations
GPTQ 4-bit
Fine-tune it
Conditional / customQLoRA10.8 GB1× RTX 3060 12GB
LoRA31.1 GB1× RTX 5090 32GB
Full fine-tune226.0 GB2× H200 141GB
QLoRA SFT on ~10k samples ≈ $8.19 (1× RTX 3060 12GB)
GLM-4 family
Elo progression across releases
API price weights · each benchmark row carries its own source badge (see methodology)