Model
Model explorer

DeepSeek-V4-Pro (Non-Think)

OPEN
DeepSeek · DeepSeek-V4 family · released Apr 24, 2026

Fast/low-latency mode of the DeepSeek-V4-Pro preview checkpoint, the default reasoning tier a user gets without selecting anything. (fast/low-latency mode; the default tier a user gets without selecting anything).

ReasoningCodingVisionFunction callingTool useAgentic
1643.5
Elo · rank #111
Parameters
1600B
Active params
49B (MoE)
Context
1M tokens
Architecture
Mixture-of-Experts (1.6T total / 49B active), hybrid Compressed Sparse Attention + Heavily Compressed Attention, Manifold-Constrained Hyper-Connections, Muon optimizer
License
MIT
Languages
API price (in/out)
$0.435 / $0.87
Modalities
text
Benchmark results
Bar shows position within the tracked field; marker = field best
GPQA DiamondReasoning72.9%#120
best: GPT-6 Astra · 96.0%
Humanity's Last ExamReasoning7.7%#108
best: Claude Opus 5 · 64.7%
IFBenchReasoning46.0%#45
best: MiniMax M3 · 83.0%
LiveCodeBenchCoding56.8%#107
best: DeepSeek-V4-Pro (Think Max) · 93.5%
MMLU-ProKnowledge82.9%#41
best: Claude Fable 5 · 91.5%
SWE-bench ProCoding52.1%#45
best: Claude Fable 5.1 · 81.2%
SWE-bench VerifiedCoding73.6%#50
best: Claude Opus 5 · 96.0%
Terminal-Bench 2.0Coding59.1%#47
best: Gemini 3.8 Flash · 89.4%
Run it locally
VRAM @ Q4
VRAM @ FP16
Fits on (Q4)
Multi-node cluster required
Throughput data unavailable.
Quantizations
Fine-tune it
Permissive
QLoRA1088.0 GB8× H200 141GB
LoRA3408.0 GBbeyond 8× B200
Full fine-tune25680.0 GBbeyond 8× B200
QLoRA SFT on ~10k samples ≈ $25.34 (8× H200 141GB)
API price $0.435/$0.87 · each benchmark row carries its own source badge (see methodology)