Model
Model explorer
Apple DCLM-Baseline 7B
OPENApple · Apple DCLM family · released Jul 1, 2024
Data-centric base LM trained on the curated DCLM-Baseline corpus (3.8T tokens + StarCoder/ProofPile2); not instruction-tuned.
ReasoningCodingVisionFunction callingTool useAgentic
62.7
Elo · rank #397
Parameters
7B
Active params
7B (dense)
Context
2K tokens
Architecture
Dense decoder-only Transformer (32 layers, 4096 hidden, 32 heads)
License
Apple Sample Code License
Languages
—
API price (in/out)
No hosted API
Modalities
text
Benchmark results
Bar shows position within the tracked field; marker = field best
best: Llama 3.1 405B · 96.9%
best: Phi-3-medium (14B) · 97.7%
best: GPT-6 Astra · 96.0%
best: Llama 3.1 405B · 96.8%
best: Claude 3 Opus · 95.4%
best: GPT-3 175B · 86.4%
best: OpenAI o3 · 92.9%
best: Claude 1 · 90.8%
best: GPT-4o mini · 93.1%
best: Palmyra Med 70B · 79.6%
best: this model · 82.9%
best: Sarvam-1 (2B) · 90.6%
best: PaLM 2 · 90.9%
Run it locally
VRAM @ Q4
—
VRAM @ FP16
14 GB
Fits on (Q4)
Multi-node cluster required
Throughput data unavailable.
Quantizations
—
Fine-tune it
Conditional / customQLoRA6.4 GB1× RTX 3060 12GB
LoRA16.6 GB1× RTX 3090 24GB
Full fine-tune114.0 GB1× H200 141GB
QLoRA SFT on ~10k samples ≈ $4.10 (1× RTX 3060 12GB)
Resources
API price weights · each benchmark row carries its own source badge (see methodology)