Model
Model explorer

gpt-oss-20b (Medium)

OPEN
OpenAI · gpt-oss family · released Aug 5, 2025

21B MoE (3.6B active) open-weight sibling that runs on consumer hardware; Apache 2.0.

ReasoningCodingVisionFunction callingTool useAgentic
1340.3
Elo · rank #163
Parameters
21B
Active params
3.6B (MoE)
Context
128K tokens
Architecture
Mixture-of-experts transformer, 32 experts/layer, top-4 active, 24 layers
License
Apache 2.0
Languages
API price (in/out)
No hosted API
Modalities
text
Benchmark results
Bar shows position within the tracked field; marker = field best
Aider PolyglotCoding26.6%#34
best: Claude Opus 4.5 · 89.4%
AIMEMath80.0%#79
best: GPT-5.2 · 100.0%
CodeforcesCoding1998#20
best: DeepSeek-V4-Pro (Think Max) · 3206
GPQA DiamondReasoning66.0%#146
best: GPT-6 Astra · 96.0%
Humanity's Last ExamReasoning7.0%#110
best: Claude Opus 5 · 64.7%
MMLUKnowledge84.0%#67
best: OpenAI o3 · 92.9%
MMMLUKnowledge73.5%#39
best: Gemini 3.1 Pro · 92.6%
SWE-bench VerifiedCoding53.2%#91
best: Claude Opus 5 · 96.0%
tau-benchAgents47.3%#23
best: Claude Opus 4.1 · 82.4%
Run it locally
VRAM @ Q4
13 GB
VRAM @ FP16
42 GB
Fits on (Q4)
RTX 4070 Ti 16GBRTX 3090 24GBRTX 4090 24GBRTX 5090 32GBM4 Pro 48GBM3 Max 128GBM3 Ultra 512GBA100 80GBH100 80GBH200 141GBB200 192GB
~225 tok/s on RTX 4090 (Q4, llama.cpp)
Quantizations
GGUF Q4 · MXFP4 (native) · GGUF Q4_K_M · GGUF Q8_0
Fine-tune it
Permissive
QLoRA15.2 GB1× RTX 4070 Ti 16GB
LoRA45.7 GB1× A100 80GB
Full fine-tune338.0 GB2× B200 192GB
QLoRA SFT on ~10k samples ≈ $1.32 (1× RTX 4070 Ti 16GB)
API price weights · each benchmark row carries its own source badge (see methodology)