Model
Model explorer

DeepSeek-V3.1 (Thinking)

OPEN
DeepSeek · DeepSeek-V3 family · released Aug 21, 2025

Thinking mode of the hybrid V3.1 checkpoint; served via deepseek-reasoner, faster than R1-0528 at similar quality

ReasoningCodingVisionFunction callingTool useAgentic
1890.4
Elo · rank #86
Parameters
671B
Active params
37B (MoE)
Context
128K tokens
Architecture
Mixture-of-Experts (671B total / 37B active) — DeepSeekMoE + Multi-head Latent Attention, hybrid think/non-think chat template
License
MIT
Languages
API price (in/out)
$0.56 / $1.68
Modalities
text
Benchmark results
Bar shows position within the tracked field; marker = field best
Aider PolyglotCoding76.3%#8
best: Claude Opus 4.5 · 89.4%
AIMEMath93.1%#30
best: GPT-5.2 · 100.0%
BrowseCompAgents30.0%#41
best: Kimi K3 · 91.2%
CodeforcesCoding2091#16
best: DeepSeek-V4-Pro (Think Max) · 3206
GPQA DiamondReasoning80.1%#85
best: GPT-6 Astra · 96.0%
Humanity's Last ExamReasoning15.9%#79
best: Claude Opus 5 · 64.7%
LiveCodeBenchCoding74.8%#59
best: DeepSeek-V4-Pro (Think Max) · 93.5%
MMLU-ProKnowledge84.8%#28
best: Claude Fable 5 · 91.5%
MMLU-ReduxKnowledge93.7%#10
best: Qwen3.7-Max · 95.0%
Run it locally
VRAM @ Q4
405 GB
VRAM @ FP16
1340 GB
Fits on (Q4)
Multi-node cluster required
Exceeds a single RTX 4090's 24GB VRAM even at Q4 (~405GB); requires multi-GPU or heavy CPU offload
Quantizations
GGUF · FP8
Fine-tune it
Permissive
QLoRA456.3 GB4× H200 141GB
LoRA1429.2 GB8× B200 192GB
Full fine-tune10769.5 GBbeyond 8× B200
QLoRA SFT on ~10k samples ≈ $19.14 (4× H200 141GB)
API price $0.56/$1.68 · each benchmark row carries its own source badge (see methodology)