Model
Model explorer

QwQ-32B-Preview

OPEN
Alibaba · QwQ family · released Nov 28, 2024

First Qwen o1-style CoT reasoning preview; superseded by QwQ-32B three months later.

ReasoningCodingVisionFunction callingTool useAgentic
1084.7
Elo · rank #207
Parameters
32.5B
Active params
32.5B (dense)
Context
32K tokens
Architecture
Dense Transformer (GQA, RoPE, SwiGLU, RMSNorm)
License
Apache 2.0
Languages
API price (in/out)
No hosted API
Modalities
text
Benchmark results
Bar shows position within the tracked field; marker = field best
AIMEMath50.0%#117
best: GPT-5.2 · 100.0%
BIG-Bench HardReasoning66.6%#68
best: ERNIE 4.5 300B-A47B · 94.3%
GPQA DiamondReasoning65.2%#152
best: GPT-6 Astra · 96.0%
IFEvalReasoning33.8%#140
best: Gemma 4 26B A4B · 98.5%
LiveCodeBenchCoding50.0%#121
best: DeepSeek-V4-Pro (Think Max) · 93.5%
MATH-500Math90.6%#44
best: GPT-5 · 99.4%
MMLU-ProKnowledge71.0%#90
best: Claude Fable 5 · 91.5%
Run it locally
VRAM @ Q4
19.85 GB
VRAM @ FP16
65 GB
Fits on (Q4)
RTX 3090 24GBRTX 4090 24GBRTX 5090 32GBM4 Pro 48GBM3 Max 128GBM3 Ultra 512GBA100 80GBH100 80GBH200 141GBB200 192GB
Throughput data unavailable.
Quantizations
GGUF
Fine-tune it
Permissive
QLoRA22.5 GB1× RTX 3090 24GB
LoRA69.6 GB1× A100 80GB
Full fine-tune522.0 GB4× H200 141GB
QLoRA SFT on ~10k samples ≈ $16.74 (1× RTX 3090 24GB)
QwQ family
Elo progression across releases
API price weights · each benchmark row carries its own source badge (see methodology)