Model
Model explorer

Mistral Small 3.1 (24B)

OPEN
Mistral AI · Mistral Small family · released Mar 17, 2025

First Mistral Small with native vision-language support; 128k context, Apache 2.0, fits a single RTX 4090 quantized.

ReasoningCodingVisionFunction callingTool useAgentic
991.2
Elo · rank #225
Parameters
24B
Active params
24B (dense)
Context
128K tokens
Architecture
Dense transformer — 24B + Pixtral-style vision encoder, Tekken tokenizer
License
Apache 2.0
Languages
API price (in/out)
$0.1 / $0.3
Modalities
text · vision
Benchmark results
Bar shows position within the tracked field; marker = field best
AI2DVision93.7%#13
best: Molmo 72B · 96.3%
ChartQAVision86.2%#18
best: MiniMax-VL-01 · 91.7%
DocVQAVision94.1%#12
best: Qwen2-VL-72B · 96.5%
GPQA DiamondReasoning46.0%#210
best: GPT-6 Astra · 96.0%
HumanEvalCoding88.4%#26
best: Claude Opus 4.5 · 99.4%
MATH-500Math69.3%#96
best: GPT-5 · 99.4%
MathVistaVision68.9%#30
best: Seed 2.1 Pro · 90.7%
MBPPCoding74.7%#26
best: Llama-3.3-Nemotron-Super-49B v1 (Reasoning On) · 91.3%
MMLU-ProKnowledge66.8%#101
best: Claude Fable 5 · 91.5%
MMLUKnowledge80.6%#84
best: OpenAI o3 · 92.9%
MMMU-ProVision49.3%#51
best: Claude Opus 4.7 · 85.5%
MMMUVision64.0%#68
best: Claude Fable 5 · 89.3%
SimpleQAKnowledge10.4%#36
best: GPT-4.5 · 62.5%
Run it locally
VRAM @ Q4
15 GB
VRAM @ FP16
55 GB
Fits on (Q4)
RTX 3090 24GBRTX 4090 24GBRTX 5090 32GBM4 Pro 48GBM3 Max 128GBM3 Ultra 512GBA100 80GBH100 80GBH200 141GBB200 192GB
Throughput data unavailable.
Quantizations
GGUF Q4 · AWQ · MLX
Fine-tune it
Permissive
QLoRA17.1 GB1× RTX 3090 24GB
LoRA51.9 GB1× A100 80GB
Full fine-tune386.0 GB4× H200 141GB
QLoRA SFT on ~10k samples ≈ $12.36 (1× RTX 3090 24GB)
API price $0.1/$0.3 · each benchmark row carries its own source badge (see methodology)