Model
Model explorer
Inkling
OPENThinking Machines Lab · Inkling family · released Jul 15, 2026
Thinking Machines Lab's (Mira Murati's startup) first public model release — a 975B/41B-active MoE with adjustable inference-time "thinking effort", calibrated uncertainty, and text/vision/audio input.
ReasoningCodingVisionFunction callingTool useAgentic
2231.4
Elo · rank #55
Parameters
975B
Active params
41B (MoE)
Context
1M tokens
Architecture
Sparse MoE, 66 layers (975B total / 41B active, 6-of-256 routed experts + 2 shared experts/token), hybrid sliding-window/global attention (5:1), relative positional features (no RoPE)
License
Apache 2.0
Languages
—
API price (in/out)
$1.87 / $4.68
Modalities
text · vision · audio
Benchmark results
Bar shows position within the tracked field; marker = field best
best: GPT-5.2 · 100.0%
best: Kimi K3 · 91.2%
best: Qwen3.8-Flash-Next · 90.6%
best: Claude Fable 5 · 1932
best: GPT-6 Astra · 96.0%
best: Claude Opus 5 · 64.7%
best: MiniMax M3 · 83.0%
best: Kimi K3 · 84.2%
best: Claude Opus 4.7 · 85.5%
best: GPT-4.5 · 62.5%
best: Claude Fable 5.1 · 81.2%
best: Claude Opus 5 · 96.0%
best: Gemini 3.8 Flash · 89.4%
Run it locally
VRAM @ Q4
600 GB
VRAM @ FP16
1900 GB
Fits on (Q4)
Multi-node cluster required
Official NVFP4 checkpoint needs Blackwell-class GPUs (~600GB VRAM); community Unsloth GGUF quants range 1-bit (~270GB) to 8-bit (~870GB) — far beyond a single RTX 4090 at any quant level.
Quantizations
NVFP4 · GGUF
Fine-tune it
PermissiveQLoRA663.0 GB4× B200 192GB
LoRA2076.8 GBbeyond 8× B200
Full fine-tune15648.8 GBbeyond 8× B200
QLoRA SFT on ~10k samples ≈ $14.66 (4× B200 192GB)
API price $1.87/$4.68 · each benchmark row carries its own source badge (see methodology)