Model
Model explorer

Kimi K2 Thinking

OPEN
Moonshot AI · Kimi K2 family · released Nov 6, 2025

Open-weight 1T MoE reasoning agent with native INT4 weights and 256K context, chaining 200-300 sequential tool calls.

ReasoningCodingVisionFunction callingTool useAgentic
2161.6
Elo · rank #64
Parameters
1000B
Active params
32B (MoE)
Context
256K tokens
Architecture
Mixture-of-Experts, 384 experts (8 active + 1 shared/token), native INT4 quantization-aware training
License
Modified MIT License
Languages
API price (in/out)
$0.6 / $2.5
Modalities
text
Benchmark results
Bar shows position within the tracked field; marker = field best
AIMEMath99.1%#6
best: GPT-5.2 · 100.0%
BrowseCompAgents60.2%#30
best: Kimi K3 · 91.2%
GPQA DiamondReasoning84.5%#64
best: GPT-6 Astra · 96.0%
Humanity's Last ExamReasoning44.9%#15
best: Claude Opus 5 · 64.7%
IFBenchReasoning68.0%#32
best: MiniMax M3 · 83.0%
LiveCodeBenchCoding83.1%#36
best: DeepSeek-V4-Pro (Think Max) · 93.5%
MMLU-ProKnowledge84.6%#30
best: Claude Fable 5 · 91.5%
MMLU-ReduxKnowledge94.4%#5
best: Qwen3.7-Max · 95.0%
SWE-bench VerifiedCoding71.3%#58
best: Claude Opus 5 · 96.0%
τ²-Bench TelecomAgents93.0%#17
best: Claude Opus 4.6 · 99.3%
Terminal-BenchCoding47.1%#2
best: Claude Opus 4.5 (High) · 59.3%
Run it locally
VRAM @ Q4
600 GB
VRAM @ FP16
2000 GB
Fits on (Q4)
Multi-node cluster required
Does not fit on a single RTX 4090 (24GB); native INT4 packs ~500-600GB combined RAM/VRAM, community reports only ~5-10 tok/s with heavy CPU/SSD offload
Quantizations
GGUF · INT4 (native QAT)
Fine-tune it
Permissive
QLoRA680.0 GB4× B200 192GB
LoRA2130.0 GBbeyond 8× B200
Full fine-tune16050.0 GBbeyond 8× B200
QLoRA SFT on ~10k samples ≈ $11.44 (4× B200 192GB)
API price $0.6/$2.5 · each benchmark row carries its own source badge (see methodology)