Model
Model explorer

OLMo 3 7B

OPEN
Allen Institute for AI (Ai2) · OLMo 3 family · released Nov 20, 2025

Lighter-weight member of the OLMo 3 family (Base/Think/Instruct) built on the Dolma 3 / Dolci stack; runs on a wider range of hardware.

ReasoningCodingVisionFunction callingTool useAgentic
843.3
Elo · rank #258
Parameters
7.3B
Active params
7.3B (dense)
Context
64K tokens
Architecture
Dense transformer (decoder-only, sliding/full attention mix)
License
Apache 2.0
Languages
1+
API price (in/out)
No hosted API
Modalities
text
Benchmark results
Bar shows position within the tracked field; marker = field best
AGIEvalReasoning64.4%#5
best: OLMo 3-Think 32B · 88.2%
AIMEMath44.3%#123
best: GPT-5.2 · 100.0%
AlpacaEvalHuman preference40.9%#26
best: Tülu 2+DPO 70B · 95.1%
BIG-Bench HardReasoning71.2%#50
best: ERNIE 4.5 300B-A47B · 94.3%
BFCL v3Agents49.8%#30
best: Hunyuan-A13B · 78.3%
GPQA DiamondReasoning40.0%#226
best: GPT-6 Astra · 96.0%
Humanity's Last ExamReasoning5.8%#120
best: Claude Opus 5 · 64.7%
HumanEval+Coding77.2%#17
best: Mistral Small 3.2 (24B) · 92.9%
IFBenchReasoning32.3%#70
best: MiniMax M3 · 83.0%
IFEvalReasoning85.6%#56
best: Gemma 4 26B A4B · 98.5%
LiveCodeBenchCoding29.5%#153
best: DeepSeek-V4-Pro (Think Max) · 93.5%
MATH-500Math87.3%#55
best: GPT-5 · 99.4%
MBPP+Coding60.2%#29
best: Llama 3.1 405B · 88.6%
MMLU-ProKnowledge52.2%#129
best: Claude Fable 5 · 91.5%
MMLUKnowledge69.1%#152
best: OpenAI o3 · 92.9%
Run it locally
VRAM @ Q4
5 GB
VRAM @ FP16
15 GB
Fits on (Q4)
RTX 3060 12GBRTX 4070 Ti 16GBRTX 3090 24GBRTX 4090 24GBRTX 5090 32GBM4 Pro 48GBM3 Max 128GBM3 Ultra 512GBA100 80GBH100 80GBH200 141GBB200 192GB
Throughput data unavailable.
Quantizations
GGUF Q4
Fine-tune it
Permissive
QLoRA6.6 GB1× RTX 3060 12GB
LoRA17.2 GB1× RTX 3090 24GB
Full fine-tune118.8 GB1× H200 141GB
QLoRA SFT on ~10k samples ≈ $4.27 (1× RTX 3060 12GB)
OLMo 3 family
Elo progression across releases
API price weights · each benchmark row carries its own source badge (see methodology)