Model
Model explorer

Qwen2.5-VL-72B

OPEN
Alibaba · Qwen2.5-VL family · released Jan 28, 2025

Gen-2.5 vision flagship with agentic GUI/computer-use grounding and 1hr+ video understanding; 3B/7B/72B family.

ReasoningCodingVisionFunction callingTool useAgentic
1171.4
Elo · rank #198
Parameters
73B
Active params
73B (dense)
Context
32K tokens
Architecture
Dense Vision-Language Transformer (window-attention ViT + Qwen2.5 LLM, mRoPE)
License
Qwen (Tongyi Qianwen) License
Languages
API price (in/out)
No hosted API
Modalities
text · vision · video
Benchmark results
Bar shows position within the tracked field; marker = field best
AI2DVision88.4%#25
best: Molmo 72B · 96.3%
ChartQAVision89.5%#6
best: MiniMax-VL-01 · 91.7%
CharXivVision49.7%#37
best: Qwen3.8-Flash-Next · 90.6%
DocVQAVision96.4%#3
best: Qwen2-VL-72B · 96.5%
GPQA DiamondReasoning49.0%#200
best: GPT-6 Astra · 96.0%
GSM8KMath95.3%#7
best: Llama 3.1 405B · 96.8%
HumanEvalCoding87.8%#32
best: Claude Opus 4.5 · 99.4%
IFEvalReasoning86.3%#48
best: Gemma 4 26B A4B · 98.5%
MATH-500Math83.0%#62
best: GPT-5 · 99.4%
MathVisionVision38.1%#18
best: Seed 2.1 Pro · 92.6%
MathVistaVision74.8%#17
best: Seed 2.1 Pro · 90.7%
MMBench (Chinese)Vision87.9%#4
best: ERNIE 4.5 VL 424B-A47B · 90.9%
MMBench (English)Vision88.0%#10
best: Qwen3.5-397B-A17B · 93.7%
MMLU-ProKnowledge71.2%#89
best: Claude Fable 5 · 91.5%
MMLU-ReduxKnowledge85.9%#29
best: Qwen3.7-Max · 95.0%
MMMU-ProVision51.1%#50
best: Claude Opus 4.7 · 85.5%
MMMUVision70.2%#53
best: Claude Fable 5 · 89.3%
OCRBenchVision885#4
best: InternVL3-78B · 906
OSWorldAgents8.8%#5
best: Seed 2.1 Pro · 78.8%
RefCOCO+Vision88.9%#5
best: DeepSeek-VL2 · 91.2%
RefCOCOVision92.7%#6
best: Qwen3.5-Omni-Plus · 95.0%
TextVQAVision83.5%#6
best: Molmo 2 8B · 85.7%
Video-MMEVision79.1%#11
best: Seed 2.1 Pro · 89.2%
Run it locally
VRAM @ Q4
42 GB
VRAM @ FP16
144 GB
Fits on (Q4)
M3 Max 128GBM3 Ultra 512GBA100 80GBH100 80GBH200 141GBB200 192GB
Throughput data unavailable.
Quantizations
GGUF · AWQ
Fine-tune it
Conditional / custom
QLoRA49.6 GB1× A100 80GB
LoRA155.5 GB1× B200 192GB
Full fine-tune1171.7 GB8× B200 192GB
QLoRA SFT on ~10k samples ≈ $51.34 (1× A100 80GB)
Qwen2.5-VL family
Elo progression across releases
API price weights · each benchmark row carries its own source badge (see methodology)