Model
Model explorer

Claude Opus 4.6

CLOSED
Anthropic · Claude 4 family · released Feb 5, 2026

First Opus-class model with a 1M-token context window in beta (Claude Developer Platform only) plus Claude Code agent teams; also exposes low/medium/high/max effort levels (reported here at default high).

ReasoningCodingVisionFunction callingTool useAgentic
2433.7
Elo · rank #39
Parameters
Undisclosed
Active params
Context
1M tokens
Architecture
Dense Transformer (undisclosed configuration)
License
Proprietary
Languages
API price (in/out)
$5 / $25
Modalities
text · vision
Benchmark results
Bar shows position within the tracked field; marker = field best
AIMEMath99.8%#4
best: GPT-5.2 · 100.0%
ARC-AGI-1Reasoning94.0%#6
best: GPT-6 Astra · 98.5%
ARC-AGI-2Reasoning68.8%#11
best: GPT-6 Astra · 95.0%
BrowseCompAgents84.0%#12
best: Kimi K3 · 91.2%
CharXivVision69.1%#31
best: Qwen3.8-Flash-Next · 90.6%
CyberGymCoding66.6%#9
best: GLM-5.3 · 84.5%
GDPval-AAAgents1606#13
best: Claude Fable 5 · 1932
GPQA DiamondReasoning91.3%#23
best: GPT-6 Astra · 96.0%
Humanity's Last ExamReasoning40.0%#30
best: Claude Opus 5 · 64.7%
LiveCodeBenchCoding84.7%#27
best: DeepSeek-V4-Pro (Think Max) · 93.5%
MMMLUKnowledge91.1%#4
best: Gemini 3.1 Pro · 92.6%
MMMU-ProVision73.9%#32
best: Claude Opus 4.7 · 85.5%
MMMUVision83.9%#12
best: Claude Fable 5 · 89.3%
OSWorld-VerifiedAgents72.7%#15
best: Claude Fable 5 · 85.0%
SWE-bench VerifiedCoding80.8%#12
best: Claude Opus 5 · 96.0%
τ²-Bench TelecomAgents99.3%#1
best: this model · 99.3%
Terminal-Bench 2.0Coding65.4%#37
best: Gemini 3.8 Flash · 89.4%
API price $5/$25 · each benchmark row carries its own source badge (see methodology)