Model
Model explorer

GPT-5.6 Sol

CLOSED
OpenAI · GPT-5.6 family · released Jul 9, 2026

Flagship (Sol) tier of the three-tier GPT-5.6 family, billed as OpenAI's best coding model yet (Terminal-Bench 2.1 SOTA); limited-preview June 26 under US govt review, GA July 9, 2026. Sibling tiers GPT-5.6 Terra and GPT-5.6 Luna are tracked as their own corpus models in the same family (OpenAI: 'durable capability tiers that can advance on their own cadence').

ReasoningCodingVisionFunction callingTool useAgentic
2915.3
Elo · rank #12
Parameters
Undisclosed
Active params
Context
1.05M tokens
Architecture
Undisclosed architecture (presumed dense transformer); flagship tier of the three durable capability tiers Sol/Terra/Luna, each tracked as its own corpus model (like Claude Opus/Sonnet/Haiku)
License
Proprietary
Languages
API price (in/out)
$5 / $30
Modalities
text · vision
Benchmark results
Bar shows position within the tracked field; marker = field best
AA-BriefcaseAgents1505#5
best: Claude Opus 5 · 1720
Agents' Last ExamAgents52.7%#2
best: GPT-6 Astra · 59.3%
ARC-AGI-1Reasoning96.5%#4
best: GPT-6 Astra · 98.5%
ARC-AGI-2Reasoning92.5%#2
best: GPT-6 Astra · 95.0%
ARC-AGI-3Reasoning7.8%#2
best: Claude Opus 5 · 30.2%
AutomationBenchAgents18.1%#9
best: Qwen3.8-Max-0902 · 50.8%
best: GPT-6 Astra · 95.9%
BrowseCompAgents90.4%#3
best: Kimi K3 · 91.2%
CursorBenchCoding67.2%#5
best: Claude Fable 5.1 · 73.4%
CyberGymCoding84.5%#2
best: GLM-5.3 · 84.5%
DeepSearchQAAgents93.0%#1
best: this model · 93.0%
DeepSWECoding72.7%#3
best: Muse Spark 1.3 · 75.4%
ExploitBenchAgents76.5%#2
best: GPT-6 Astra · 100.0%
FrontierBenchAgents34.4%#2
best: Claude Opus 5 · 43.3%
best: Claude Fable 5 · 64.9%
FrontierCodeCoding47.5%#5
best: Claude Fable 5 · 53.5%
best: this model · 89.0%
best: this model · 83.0%
GDP.PDFVision40.0%#1
best: this model · 40.0%
GDPval-AAAgents1747.8#8
best: Claude Fable 5 · 1932
GPQA DiamondReasoning94.6%#2
best: GPT-6 Astra · 96.0%
best: Grok 4.6 · 15.8%
HLE-VerifiedKnowledge54.5%#2
best: Gemini 3.8 Flash · 54.9%
Humanity's Last ExamReasoning47.2%#12
best: Claude Opus 5 · 64.7%
IFBenchReasoning72.7%#25
best: MiniMax M3 · 83.0%
JobBenchAgents45.4%#6
best: Claude Opus 5 · 65.7%
LVBenchVision82.1%#3
best: Gemini 3.8 Flash · 87.8%
MMMU-ProVision83.0%#4
best: Claude Opus 4.7 · 85.5%
best: Muse Spark 1.3 · 98.1%
best: GPT-6 Astra · 100.0%
OSWorld 2.0Agents62.6%#5
best: Claude Fable 5.1 · 77.9%
ScreenSpot-ProVision76.9%#3
best: GPT-6 Astra · 92.7%
SRE-BenchCoding55.9%#2
best: GPT-6 Astra · 88.0%
best: Muse Spark 1.3 · 59.4%
SWE-bench ProCoding64.6%#7
best: Claude Fable 5.1 · 81.2%
τ³-BankingAgents33.0%#1
best: this model · 33.0%
Terminal-Bench 2.0Coding88.8%#2
best: Gemini 3.8 Flash · 89.4%
Terminal-Bench 3.0Agents34.6%#1
best: this model · 34.6%
Terminal-Bench 4.0Agents37.3%#6
best: Claude Mythos 5.1 · 60.9%
best: GPT-6 Astra · 64.6%
best: GLM-5.3-Flash · 78.4%
Vals Finance AgentAgents53.8%#4
best: Gemini 3.8 Flash · 61.4%
Run it locally
Closed weights — available via API only. No local deployment.
Input / M tok
$5
Output / M tok
$30
API price $5/$30 · each benchmark row carries its own source badge (see methodology)