Model
Model explorer

Inkling

OPEN
Thinking Machines Lab · Inkling family · released Jul 15, 2026

Thinking Machines Lab's (Mira Murati's startup) first public model release — a 975B/41B-active MoE with adjustable inference-time "thinking effort", calibrated uncertainty, and text/vision/audio input.

ReasoningCodingVisionFunction callingTool useAgentic
2231.4
Elo · rank #55
Parameters
975B
Active params
41B (MoE)
Context
1M tokens
Architecture
Sparse MoE, 66 layers (975B total / 41B active, 6-of-256 routed experts + 2 shared experts/token), hybrid sliding-window/global attention (5:1), relative positional features (no RoPE)
License
Apache 2.0
Languages
API price (in/out)
$1.87 / $4.68
Modalities
text · vision · audio
Benchmark results
Bar shows position within the tracked field; marker = field best
AIMEMath97.1%#7
best: GPT-5.2 · 100.0%
BrowseCompAgents77.1%#21
best: Kimi K3 · 91.2%
CharXivVision78.1%#21
best: Qwen3.8-Flash-Next · 90.6%
GDPval-AAAgents1238#35
best: Claude Fable 5 · 1932
GPQA DiamondReasoning87.2%#49
best: GPT-6 Astra · 96.0%
Humanity's Last ExamReasoning29.7%#52
best: Claude Opus 5 · 64.7%
IFBenchReasoning79.8%#8
best: MiniMax M3 · 83.0%
MCP AtlasAgents74.1%#5
best: Kimi K3 · 84.2%
MMMU-ProVision73.5%#35
best: Claude Opus 4.7 · 85.5%
SimpleQAKnowledge43.9%#7
best: GPT-4.5 · 62.5%
SWE-bench ProCoding54.3%#38
best: Claude Fable 5.1 · 81.2%
SWE-bench VerifiedCoding77.6%#31
best: Claude Opus 5 · 96.0%
Terminal-Bench 2.0Coding63.8%#38
best: Gemini 3.8 Flash · 89.4%
Run it locally
VRAM @ Q4
600 GB
VRAM @ FP16
1900 GB
Fits on (Q4)
Multi-node cluster required
Official NVFP4 checkpoint needs Blackwell-class GPUs (~600GB VRAM); community Unsloth GGUF quants range 1-bit (~270GB) to 8-bit (~870GB) — far beyond a single RTX 4090 at any quant level.
Quantizations
NVFP4 · GGUF
Fine-tune it
Permissive
QLoRA663.0 GB4× B200 192GB
LoRA2076.8 GBbeyond 8× B200
Full fine-tune15648.8 GBbeyond 8× B200
QLoRA SFT on ~10k samples ≈ $14.66 (4× B200 192GB)
Inkling family
Elo progression across releases
API price $1.87/$4.68 · each benchmark row carries its own source badge (see methodology)