Model
Model explorer

GPT-6 Astra

CLOSED
OpenAI · GPT-6 family · released Sep 3, 2026

OpenAI's first GPT-6-generation model, launched September 3, 2026 as `gpt-6-astra` after a month of staged safety disclosure: previewed August 1 through ten solved open mathematics problems, then confirmed on September 1 as the first model OpenAI has ever classified **Critical** for cybersecurity under its Preparedness Framework (it can find and weaponise zero-days in hardened targets without human direction). Rolling out first to enterprises in the Trusted Access Program, with API, Plus, Pro, Business and Enterprise access following; advanced offensive-cyber capability stays gated behind Daybreak Blue. $10/$50 per Mtok (cached input $1.00, cache writes $12.50; >272K-token prompts billed at 2x input / 1.5x output; Fast mode 2x at up to 2.5x speed), 1,050,000-token context, 128K max output, April 30 2026 knowledge cutoff. PROVENANCE CAVEATS — every recorded row is OpenAI-run at maximum effort, and OpenAI does not say whether standard users can select that effort. Two headline claims are deliberately NOT recorded here: the ARC-AGI-3 99.9% (run through a Responses-API harness that retains reasoning across turns and compacts context — OpenAI has itself shown those settings can triple ARC-AGI-3 scores without changing the model, so the number measures the agent system; the public ARC Prize leaderboard's best result is Claude Opus 5 at ~30.2%, and the viral 98.6%/99.9% figures were flagged as unverified by third-party coverage) and the FrontierMath Tier 4 98% saturation claim (no independent corroboration). Its DeepSWE lead is also narrower than the launch chart suggests: OpenAI's own chart omits Muse Spark 1.3, which reports 75.4% against Astra's 74.1%.

ReasoningCodingVisionFunction callingTool useAgentic
3305.4
Elo · rank #1
Parameters
Undisclosed
Active params
Context
1.05M tokens
Architecture
Proprietary frontier reasoning model (architecture and size undisclosed); 1.05M-token context, text+image in / text out, five selectable reasoning efforts (low → max)
License
Proprietary
Languages
API price (in/out)
$10 / $50
Modalities
text · vision
Benchmark results
Bar shows position within the tracked field; marker = field best
Agents' Last ExamAgents59.3%#1
best: this model · 59.3%
ARC-AGI-1Reasoning98.5%#1
best: this model · 98.5%
ARC-AGI-2Reasoning95.0%#1
best: this model · 95.0%
AutomationBenchAgents41.4%#5
best: Qwen3.8-Max-0902 · 50.8%
best: this model · 95.9%
ExploitBenchAgents100.0%#1
best: this model · 100.0%
best: Claude Fable 5 · 64.9%
FrontierCodeCoding53.3%#3
best: Claude Fable 5 · 53.5%
GPQA DiamondReasoning96.0%#1
best: this model · 96.0%
best: Muse Spark 1.3 · 98.1%
best: this model · 100.0%
ScreenSpot-ProVision92.7%#1
best: this model · 92.7%
SRE-BenchCoding88.0%#1
best: this model · 88.0%
Terminal-Bench 4.0Agents57.7%#2
best: Claude Mythos 5.1 · 60.9%
best: this model · 64.6%
Run it locally
Closed weights — available via API only. No local deployment.
Input / M tok
$10
Output / M tok
$50
GPT-6 family
Elo progression across releases
API price $10/$50 · each benchmark row carries its own source badge (see methodology)