Model
Model explorer
Qwen2-Math-72B
OPENAlibaba · Qwen2-Math family · released Aug 8, 2024
First Qwen math specialist; Instruct variant beat GPT-4o and Claude-3.5-Sonnet on math benchmarks.
ReasoningCodingVisionFunction callingTool useAgentic
1145.4
Elo · rank #201
Parameters
72.7B
Active params
72.7B (dense)
Context
4K tokens
Architecture
Dense decoder-only Transformer (GQA), math-specialized fine-tune of Qwen2-72B
License
Tongyi Qianwen License Agreement
Languages
—
API price (in/out)
No hosted API
Modalities
text
Benchmark results
Bar shows position within the tracked field; marker = field best
best: GPT-5.2 · 100.0%
best: Falcon-H1 34B · 69.4%
best: Doubao-Seed-1.6 · 96.0%
best: Llama 3.1 405B · 96.8%
best: GPT-5 · 99.4%
best: Falcon-H1 34B · 83.6%
best: Doubao-1.5-Pro · 59.8%
Run it locally
VRAM @ Q4
42 GB
VRAM @ FP16
145 GB
Fits on (Q4)
M3 Max 128GBM3 Ultra 512GBA100 80GBH100 80GBH200 141GBB200 192GB
72B dense weights exceed a single RTX 4090's 24GB VRAM even at Q4.
Quantizations
GGUF · AWQ · GPTQ
Fine-tune it
Conditional / customQLoRA49.4 GB1× A100 80GB
LoRA154.9 GB1× B200 192GB
Full fine-tune1166.8 GB8× B200 192GB
QLoRA SFT on ~10k samples ≈ $51.13 (1× A100 80GB)
API price weights · each benchmark row carries its own source badge (see methodology)