Model
Model explorer

DeepSeek-V3.1 (Non-Thinking)

OPEN
DeepSeek · DeepSeek-V3 family · released Aug 21, 2025

Non-thinking mode of the hybrid V3.1 checkpoint; served via deepseek-chat, now deprecated in favor of V3.2/V4

ReasoningCodingVisionFunction callingTool useAgentic
1517.5
Elo · rank #135
Parameters
671B
Active params
37B (MoE)
Context
128K tokens
Architecture
Mixture-of-Experts (671B total / 37B active) — DeepSeekMoE + Multi-head Latent Attention, hybrid think/non-think chat template
License
MIT
Languages
API price (in/out)
$0.56 / $1.68
Modalities
text
Benchmark results
Bar shows position within the tracked field; marker = field best
Aider PolyglotCoding68.4%#15
best: Claude Opus 4.5 · 89.4%
AIMEMath66.3%#109
best: GPT-5.2 · 100.0%
Arena EloHuman preference1418#16
best: Claude Fable 5 · 1505
GPQA DiamondReasoning74.9%#112
best: GPT-6 Astra · 96.0%
Humanity's Last ExamReasoning6.0%#118
best: Claude Opus 5 · 64.7%
IFBenchReasoning38.0%#58
best: MiniMax M3 · 83.0%
LiveCodeBenchCoding56.4%#110
best: DeepSeek-V4-Pro (Think Max) · 93.5%
MMLU-ProKnowledge83.7%#35
best: Claude Fable 5 · 91.5%
MMLU-ReduxKnowledge91.8%#19
best: Qwen3.7-Max · 95.0%
SWE-bench VerifiedCoding66.0%#71
best: Claude Opus 5 · 96.0%
Terminal-BenchCoding31.3%#11
best: Claude Opus 4.5 (High) · 59.3%
Run it locally
VRAM @ Q4
405 GB
VRAM @ FP16
1340 GB
Fits on (Q4)
Multi-node cluster required
Exceeds a single RTX 4090's 24GB VRAM even at Q4 (~405GB); requires multi-GPU or heavy CPU offload
Quantizations
GGUF · FP8
Fine-tune it
Permissive
QLoRA456.3 GB4× H200 141GB
LoRA1429.2 GB8× B200 192GB
Full fine-tune10769.5 GBbeyond 8× B200
QLoRA SFT on ~10k samples ≈ $19.14 (4× H200 141GB)
API price $0.56/$1.68 · each benchmark row carries its own source badge (see methodology)