Model
Model explorer

MPT-7B-StoryWriter-65k+

OPEN
MosaicML · MPT family · released May 5, 2023

Fiction-focused MPT-7B fine-tune trained at 65k tokens via ALiBi; demonstrated generations to ~84k tokens on a single A100-80GB node.

ReasoningCodingVisionFunction callingTool useAgentic
-474.3
Elo · rank #458
Parameters
6.7B
Active params
6.7B (dense)
Context
65K tokens
Architecture
Decoder-only Transformer w/ ALiBi (65K-token fiction fine-tune)
License
Apache-2.0
Languages
API price (in/out)
No hosted API
Modalities
text
Benchmark results
Bar shows position within the tracked field; marker = field best
ARC-ChallengeReasoning45.6%#135
best: Llama 3.1 405B · 96.9%
DROPReasoning0.32#78
best: Hunyuan-T1 · 93.1
GSM8KMath0.0%#202
best: Llama 3.1 405B · 96.8%
HellaSwagReasoning74.1%#117
best: Claude 3 Opus · 95.4%
MMLUKnowledge28.8%#272
best: OpenAI o3 · 92.9%
TruthfulQAKnowledge36.1%#90
best: Phi-3.5-MoE (16x3.8B, 6.6B active) · 77.5%
WinoGrandeReasoning51.1%#138
best: PaLM 2 · 90.9%
Run it locally
VRAM @ Q4
4.4 GB
VRAM @ FP16
13.4 GB
Fits on (Q4)
RTX 3060 12GBRTX 4070 Ti 16GBRTX 3090 24GBRTX 4090 24GBRTX 5090 32GBM4 Pro 48GBM3 Max 128GBM3 Ultra 512GBA100 80GBH100 80GBH200 141GBB200 192GB
Throughput data unavailable.
Quantizations
GGUF Q4 · GGML
Fine-tune it
Permissive
QLoRA6.2 GB1× RTX 3060 12GB
LoRA15.9 GB1× RTX 4070 Ti 16GB
Full fine-tune109.2 GB1× H200 141GB
QLoRA SFT on ~10k samples ≈ $3.92 (1× RTX 3060 12GB)
API price weights · each benchmark row carries its own source badge (see methodology)