Model
Model explorer

OpenAI o4-mini

CLOSED
OpenAI · o4 family · released Apr 16, 2025

Compact reasoning model with strong price/performance and native tool use.

ReasoningCodingVisionFunction callingTool useAgentic
1852.4
Elo · rank #91
Parameters
Undisclosed
Active params
Context
200K tokens
Architecture
Dense transformer reasoning model, compact (exact size undisclosed)
License
Proprietary (OpenAI API Terms)
Languages
API price (in/out)
$1.1 / $4.4
Modalities
text · vision
Benchmark results
Bar shows position within the tracked field; marker = field best
Aider PolyglotCoding72.0%#12
best: Claude Opus 4.5 · 89.4%
AIMEMath93.4%#26
best: GPT-5.2 · 100.0%
BrowseCompAgents28.3%#42
best: Kimi K3 · 91.2%
CharXivVision72.0%#29
best: Qwen3.8-Flash-Next · 90.6%
CodeforcesCoding2719#6
best: DeepSeek-V4-Pro (Think Max) · 3206
DROPReasoning77.7#34
best: Hunyuan-T1 · 93.1
GPQA DiamondReasoning81.4%#78
best: GPT-6 Astra · 96.0%
Humanity's Last ExamReasoning17.7%#75
best: Claude Opus 5 · 64.7%
HumanEvalCoding97.3%#3
best: Claude Opus 4.5 · 99.4%
LiveCodeBenchCoding82.2%#40
best: DeepSeek-V4-Pro (Think Max) · 93.5%
MATH-500Math97.5%#13
best: GPT-5 · 99.4%
MathVistaVision84.3%#10
best: Seed 2.1 Pro · 90.7%
MGSMMath93.7%#1
best: this model · 93.7%
MMLUKnowledge90.0%#12
best: OpenAI o3 · 92.9%
MMMUVision81.6%#20
best: Claude Fable 5 · 89.3%
PutnamBenchMath0.3%#3
best: DeepSeek-Prover-V2 671B · 7.1%
SimpleQAKnowledge20.2%#24
best: GPT-4.5 · 62.5%
SWE-bench VerifiedCoding68.1%#67
best: Claude Opus 5 · 96.0%
Run it locally
Closed weights — available via API only. No local deployment.
Input / M tok
$1.1
Output / M tok
$4.4
o4 family
Elo progression across releases
API price $1.1/$4.4 · each benchmark row carries its own source badge (see methodology)