Model
Model explorer

Claude Opus 4.5

CLOSED
Anthropic · Claude 4 family · released Nov 24, 2025

Efficiency-focused flagship: 67% price cut vs Opus 4.1 and state-of-the-art 80.9% SWE-bench Verified at launch.

ReasoningCodingVisionFunction callingTool useAgentic
2228.2
Elo · rank #56
Parameters
Undisclosed
Active params
Context
200K tokens
Architecture
Transformer (architecture undisclosed)
License
Proprietary
Languages
API price (in/out)
$5 / $25
Modalities
text · vision
Benchmark results
Bar shows position within the tracked field; marker = field best
Aider PolyglotCoding89.4%#1
best: this model · 89.4%
AIMEMath93.0%#31
best: GPT-5.2 · 100.0%
ARC-AGI-1Reasoning80.0%#9
best: GPT-6 Astra · 98.5%
ARC-AGI-2Reasoning37.6%#16
best: GPT-6 Astra · 95.0%
BrowseCompAgents67.8%#27
best: Kimi K3 · 91.2%
CyberGymCoding50.6%#11
best: GLM-5.3 · 84.5%
GDPval-AAAgents1416#27
best: Claude Fable 5 · 1932
GPQA DiamondReasoning87.0%#50
best: GPT-6 Astra · 96.0%
Humanity's Last ExamReasoning43.2%#19
best: Claude Opus 5 · 64.7%
HumanEvalCoding99.4%#1
best: this model · 99.4%
LiveCodeBenchCoding83.7%#32
best: DeepSeek-V4-Pro (Think Max) · 93.5%
MMLUKnowledge90.8%#4
best: OpenAI o3 · 92.9%
MMMU-ProVision70.6%#36
best: Claude Opus 4.7 · 85.5%
MMMUVision80.7%#22
best: Claude Fable 5 · 89.3%
OSWorld-VerifiedAgents66.3%#18
best: Claude Fable 5 · 85.0%
SWE-bench ProCoding52.0%#46
best: Claude Fable 5.1 · 81.2%
SWE-bench VerifiedCoding80.9%#10
best: Claude Opus 5 · 96.0%
τ²-Bench TelecomAgents98.2%#5
best: Claude Opus 4.6 · 99.3%
Terminal-Bench 2.0Coding59.8%#43
best: Gemini 3.8 Flash · 89.4%
API price $5/$25 · each benchmark row carries its own source badge (see methodology)