Model
Model explorer
Step 3.7 Flash
OPENStepFun · Step family · released May 29, 2026
Newest StepFun release (successor to Step 3.5 Flash), adding native multimodal (image/GUI/document) support that the text-only Step 3.5 Flash lacked. Not previously in corpus (corpus's newest StepFun entry is Step-3 from 2025-07-31). No shipped 'Step 4' flagship found -- only reported as in training.
ReasoningCodingVisionFunction callingTool useAgentic
2212.7
Elo · rank #61
Parameters
198B
Active params
11B (MoE)
Context
262K tokens
Architecture
Sparse Mixture-of-Experts with 1.8B vision encoder + 196B language backbone, 198B total / 11B active
License
Apache 2.0
Languages
—
API price (in/out)
$0.2 / $1.15
Modalities
text · vision
Benchmark results
Bar shows position within the tracked field; marker = field best
best: GPT-5.2 · 100.0%
best: Kimi K3 · 91.2%
best: Claude Fable 5 · 1932
best: GPT-6 Astra · 96.0%
best: Claude Opus 5 · 64.7%
best: Claude Opus 4.7 · 85.5%
best: Claude Fable 5.1 · 81.2%
best: Claude Opus 5 · 96.0%
best: Gemini 3.8 Flash · 89.4%
Run it locally
VRAM @ Q4
—
VRAM @ FP16
—
Fits on (Q4)
Multi-node cluster required
Throughput data unavailable.
Quantizations
—
Fine-tune it
PermissiveQLoRA134.6 GB1× H200 141GB
LoRA421.7 GB4× H200 141GB
Full fine-tune3177.9 GBbeyond 8× B200
QLoRA SFT on ~10k samples ≈ $5.69 (1× H200 141GB)
API price $0.2/$1.15 · each benchmark row carries its own source badge (see methodology)