Model
Model explorer

Gemini 3.8 Flash

CLOSED
Google · Gemini 3.8 family · released Sep 2, 2026

Google's September 2, 2026 Flash — its third in 43 days — at identical specifications and identical introductory pricing to 3.7 Flash ($0.75/$3.75 per Mtok until December 31, 2026, then $1.50/$7.50), and ahead of 3.7 Flash on every published row, so there is no reason to stay on the older model. Google's pitch is that it 'works harder': more reasoning steps and iterative tool calls on complex tasks. Available in AI Studio, the Gemini API, Antigravity, Gemini Enterprise and the Gemini app for AI Pro/Ultra subscribers. The honest reading of Google's own 14-row table is that it wins 8 and loses 5 to Claude Opus 5, and the losses are where the work is open-ended: Terminal-Bench 4.0 19.1 vs 51.8, OSWorld-2.0 59.0 vs 75.4, GDPval-AA v2 1545 vs 1824. It is excellent at bounded professional tasks — finance, law, charts, video, biology — at roughly a seventh of Opus 5's price, and its DeepSWE v1.1 73.7 is within noise of Opus 5's 74.0. Widely-circulated DeepSWE figures of 74% for this model come from the public Datacurve leaderboard rather than Google's table. Its FrontierCode 1.1 rows are deliberately omitted here: the only published values come from a competitor's comparison table and are identical to 3.7 Flash's, which cannot be distinguished from a carried-forward figure.

ReasoningCodingVisionFunction callingTool useAgentic
2949.4
Elo · rank #10
Parameters
Undisclosed
Active params
Undisclosed
Context
1M tokens
Architecture
Sparse Mixture-of-Experts Transformer (configuration undisclosed); natively multimodal (text/image/video/audio/PDF in, text out), 1,048,576-token input / 65,536-token output, thinking at low/medium/high
License
Proprietary
Languages
API price (in/out)
$0.75 / $3.75
Modalities
text · vision · audio · video
Benchmark results
Bar shows position within the tracked field; marker = field best
CharXivVision86.2%#4
best: Qwen3.8-Flash-Next · 90.6%
DeepSWECoding73.7%#2
best: Muse Spark 1.3 · 75.4%
GDP.PDFVision35.0%#3
best: GPT-5.6 Sol · 40.0%
GDPval-AAAgents1545#18
best: Claude Fable 5 · 1932
best: Grok 4.6 · 15.8%
HLE-VerifiedKnowledge54.9%#1
best: this model · 54.9%
LVBenchVision87.8%#1
best: this model · 87.8%
OSWorld 2.0Agents59.0%#6
best: Claude Fable 5.1 · 77.9%
Terminal-Bench 2.0Coding89.4%#1
best: this model · 89.4%
Terminal-Bench 4.0Agents19.1%#7
best: Claude Mythos 5.1 · 60.9%
Vals Finance AgentAgents61.4%#1
best: this model · 61.4%
Run it locally
Closed weights — available via API only. No local deployment.
Input / M tok
$0.75
Output / M tok
$3.75
Gemini 3.8 family
Elo progression across releases
API price $0.75/$3.75 · each benchmark row carries its own source badge (see methodology)