Model
Model explorer

Llama 4 Maverick

OPEN
Meta · Llama 4 family · released Apr 5, 2025

Natively multimodal MoE flagship; 17B active over 400B total across 128 experts, 1M-token context.

ReasoningCodingVisionFunction callingTool useAgentic
1216.1
Elo · rank #187
Parameters
400B
Active params
17B (MoE)
Context
1M tokens
Architecture
MoE, 128 experts (routed + shared) with early-fusion native multimodality, 17B active / 400B total
License
Llama 4 Community License Agreement
Languages
12+
API price (in/out)
No hosted API
Modalities
text · vision
Benchmark results
Bar shows position within the tracked field; marker = field best
ChartQAVision85.3%#24
best: MiniMax-VL-01 · 91.7%
DocVQAVision91.6%#28
best: Qwen2-VL-72B · 96.5%
GPQA DiamondReasoning69.8%#132
best: GPT-6 Astra · 96.0%
LiveCodeBenchCoding43.4%#130
best: DeepSeek-V4-Pro (Think Max) · 93.5%
MATH-500Math61.2%#110
best: GPT-5 · 99.4%
MathVistaVision73.7%#18
best: Seed 2.1 Pro · 90.7%
MBPPCoding77.6%#21
best: Llama-3.3-Nemotron-Super-49B v1 (Reasoning On) · 91.3%
MGSMMath92.3%#3
best: OpenAI o4-mini · 93.7%
MMLU-ProKnowledge80.5%#55
best: Claude Fable 5 · 91.5%
MMLUKnowledge85.5%#55
best: OpenAI o3 · 92.9%
MMMU-ProVision59.6%#44
best: Claude Opus 4.7 · 85.5%
MMMUVision73.4%#40
best: Claude Fable 5 · 89.3%
Run it locally
VRAM @ Q4
210 GB
VRAM @ FP16
800 GB
Fits on (Q4)
M3 Ultra 512GB
400B total MoE weights (~210GB at Q4) far exceed a single RTX 4090; needs multi-GPU/datacenter hardware — no single-4090 benchmark exists.
Quantizations
GGUF · AWQ · MLX
Fine-tune it
Conditional / custom
QLoRA272.0 GB2× H200 141GB
LoRA852.0 GB8× H200 141GB
Full fine-tune6420.0 GBbeyond 8× B200
QLoRA SFT on ~10k samples ≈ $8.79 (2× H200 141GB)
Llama 4 family
Elo progression across releases
API price weights · each benchmark row carries its own source badge (see methodology)