Model
Model explorer

OpenAI o3-mini (High)

CLOSED
OpenAI · o3 family · released Jan 31, 2025

Highest reasoning-effort setting: best accuracy of the o3-mini family, higher latency/cost per query.

ReasoningCodingVisionFunction callingTool useAgentic
1622.7
Elo · rank #113
Parameters
Undisclosed
Active params
Context
200K tokens
Architecture
Dense transformer reasoning model (size undisclosed)
License
Proprietary
Languages
API price (in/out)
$1.1 / $4.4
Modalities
text
Benchmark results
Bar shows position within the tracked field; marker = field best
AIMEMath87.3%#57
best: GPT-5.2 · 100.0%
CodeforcesCoding2130#13
best: DeepSeek-V4-Pro (Think Max) · 3206
DROPReasoning80.6#24
best: Hunyuan-T1 · 93.1
GDPval-AAAgents468#50
best: Claude Fable 5 · 1932
GPQA DiamondReasoning79.7%#89
best: GPT-6 Astra · 96.0%
Humanity's Last ExamReasoning12.0%#86
best: Claude Opus 5 · 64.7%
HumanEvalCoding97.6%#2
best: Claude Opus 4.5 · 99.4%
IFBenchReasoning67.0%#34
best: MiniMax M3 · 83.0%
LiveCodeBenchCoding67.4%#77
best: DeepSeek-V4-Pro (Think Max) · 93.5%
MATH-500Math97.9%#10
best: GPT-5 · 99.4%
MGSMMath92.0%#4
best: OpenAI o4-mini · 93.7%
MMLUKnowledge86.9%#37
best: OpenAI o3 · 92.9%
SimpleQAKnowledge13.8%#28
best: GPT-4.5 · 62.5%
SWE-bench VerifiedCoding49.3%#94
best: Claude Opus 5 · 96.0%
Terminal-Bench 2.0Coding4.0%#77
best: Gemini 3.8 Flash · 89.4%
Run it locally
Closed weights — available via API only. No local deployment.
Input / M tok
$1.1
Output / M tok
$4.4
API price $1.1/$4.4 · each benchmark row carries its own source badge (see methodology)