Model
Model explorer

Dolphin 2.5 Mixtral 8x7B

OPEN

QLoRA fine-tune of Mixtral 8x7B (Apache-2.0) with refusals filtered out and Dolphin-Coder/MagiCoder data added; a benchmark of the late-2023 'uncensored' MoE fine-tune wave.

ReasoningCodingVisionFunction callingTool useAgentic
Elo · unrated
Parameters
46.7B
Active params
12.9B (MoE)
Context
16K tokens
Architecture
Sparse Mixture-of-Experts (8 experts, top-2 routing; Mixtral 8x7B base)
License
Apache-2.0
Languages
API price (in/out)
No hosted API
Modalities
text
Benchmark results
Bar shows position within the tracked field; marker = field best
Run it locally
VRAM @ Q4
26 GB
VRAM @ FP16
93 GB
Fits on (Q4)
RTX 5090 32GBM4 Pro 48GBM3 Max 128GBM3 Ultra 512GBA100 80GBH100 80GBH200 141GBB200 192GB
Q4 (~26GB) slightly exceeds a single RTX 4090's 24GB VRAM; typically run with partial CPU offload or reduced context.
Quantizations
GGUF Q4 · GGUF Q5 · GGUF Q8 · AWQ · GPTQ
Fine-tune it
Permissive
QLoRA31.8 GB1× RTX 5090 32GB
LoRA99.5 GB1× H200 141GB
Full fine-tune749.5 GB4× B200 192GB
QLoRA SFT on ~10k samples ≈ $6.29 (1× RTX 5090 32GB)
Dolphin family
Elo progression across releases
API price weights · each benchmark row carries its own source badge (see methodology)