Model
Model explorer
Skywork-MoE-Base
OPENKunlun Tech · Skywork MoE family · released Jun 3, 2024
146B/22B-active MoE upcycled from Skywork-13B; first big open MoE runnable on a single 8x4090 server via FP8
ReasoningCodingVisionFunction callingTool useAgentic
491.9
Elo · rank #314
Parameters
146B
Active params
22B (MoE)
Context
8K tokens
Architecture
Mixture-of-Experts, 16 experts, 22B active (upcycled from Skywork-13B dense checkpoints)
License
Skywork Community License
Languages
2+
API price (in/out)
No hosted API
Modalities
text
Benchmark results
Bar shows position within the tracked field; marker = field best
Run it locally
VRAM @ Q4
—
VRAM @ FP16
292 GB
Fits on (Q4)
Multi-node cluster required
Official FP8 checkpoint reaches ~2,200 tok/s aggregate across an 8x RTX4090 cluster via non-uniform tensor parallelism -- not a single-GPU llama.cpp Q4 measurement
Quantizations
FP8
Fine-tune it
Conditional / customQLoRA99.3 GB1× H200 141GB
LoRA311.0 GB2× B200 192GB
Full fine-tune2343.3 GBbeyond 8× B200
QLoRA SFT on ~10k samples ≈ $11.38 (1× H200 141GB)
API price weights · each benchmark row carries its own source badge (see methodology)