Model
Model explorer

Seed-OSS-36B

OPEN
ByteDance · Seed family · released Aug 20, 2025

Apache-2.0 open release with native 512K context and a tunable thinking-length budget; trained on only 12T tokens.

ReasoningCodingVisionFunction callingTool useAgentic
1576.9
Elo · rank #122
Parameters
36B
Active params
36B (dense)
Context
512K tokens
Architecture
Dense Transformer (RoPE, GQA, RMSNorm, SwiGLU)
License
Apache-2.0
Languages
API price (in/out)
No hosted API
Modalities
text
Benchmark results
Bar shows position within the tracked field; marker = field best
AIMEMath91.7%#40
best: GPT-5.2 · 100.0%
ARC-AGI-2Reasoning1.4%#29
best: GPT-6 Astra · 95.0%
GPQA DiamondReasoning71.4%#125
best: GPT-6 Astra · 96.0%
Humanity's Last ExamReasoning10.1%#97
best: Claude Opus 5 · 64.7%
IFEvalReasoning85.8%#53
best: Gemma 4 26B A4B · 98.5%
LiveCodeBenchCoding67.4%#78
best: DeepSeek-V4-Pro (Think Max) · 93.5%
MMLU-ProKnowledge82.7%#43
best: Claude Fable 5 · 91.5%
MMLUKnowledge87.4%#34
best: OpenAI o3 · 92.9%
MMMLUKnowledge78.4%#31
best: Gemini 3.1 Pro · 92.6%
SimpleQAKnowledge9.7%#38
best: GPT-4.5 · 62.5%
SuperGPQAReasoning55.7%#17
best: Qwen3.7-Max · 73.6%
SWE-bench VerifiedCoding56.0%#89
best: Claude Opus 5 · 96.0%
tau-benchAgents70.4%#6
best: Claude Opus 4.1 · 82.4%
Run it locally
VRAM @ Q4
21.8 GB
VRAM @ FP16
72 GB
Fits on (Q4)
RTX 3090 24GBRTX 4090 24GBRTX 5090 32GBM4 Pro 48GBM3 Max 128GBM3 Ultra 512GBA100 80GBH100 80GBH200 141GBB200 192GB
Throughput data unavailable.
Quantizations
GGUF Q4 · GGUF Q8
Fine-tune it
Permissive
QLoRA24.7 GB1× RTX 5090 32GB
LoRA76.9 GB1× A100 80GB
Full fine-tune578.0 GB4× B200 192GB
QLoRA SFT on ~10k samples ≈ $17.55 (1× RTX 5090 32GB)
Seed family
Elo progression across releases
API price weights · each benchmark row carries its own source badge (see methodology)