Model
Model explorer

H2O-Danube 1.8B

OPEN
H2O.ai · H2O-Danube family · released Jan 30, 2024

H2O.ai's debut small LLM, Llama-style architecture, tuned for efficient on-device/mobile inference.

ReasoningCodingVisionFunction callingTool useAgentic
-339.0
Elo · rank #450
Parameters
1.8B
Active params
1.8B (dense)
Context
16K tokens
Architecture
Dense Llama-2/Mistral-style transformer with sliding-window attention + GQA
License
Apache 2.0
Languages
1+
API price (in/out)
No hosted API
Modalities
text
Benchmark results
Bar shows position within the tracked field; marker = field best
ARC-ChallengeReasoning39.3%#150
best: Llama 3.1 405B · 96.9%
ARC-EasyReasoning67.5%#54
best: Phi-3-medium (14B) · 97.7%
GSM8KMath15.5%#169
best: Llama 3.1 405B · 96.8%
HellaSwagReasoning67.6%#133
best: Claude 3 Opus · 95.4%
MMLUKnowledge33.5%#264
best: OpenAI o3 · 92.9%
MT-BenchHuman preference5.52#66
best: Hunyuan-Large (A52B) · 9.4
OpenBookQAReasoning39.2%#45
best: Claude 1 · 90.8%
PIQAReasoning76.7%#64
best: GPT-4o mini · 93.1%
TriviaQAKnowledge36.3%#47
best: Sarvam-1 (2B) · 90.6%
TruthfulQAKnowledge40.8%#78
best: Phi-3.5-MoE (16x3.8B, 6.6B active) · 77.5%
WinoGrandeReasoning65.3%#121
best: PaLM 2 · 90.9%
Run it locally
VRAM @ Q4
1.4 GB
VRAM @ FP16
3.8 GB
Fits on (Q4)
RTX 3060 12GBRTX 4070 Ti 16GBRTX 3090 24GBRTX 4090 24GBRTX 5090 32GBM4 Pro 48GBM3 Max 128GBM3 Ultra 512GBA100 80GBH100 80GBH200 141GBB200 192GB
Throughput data unavailable.
Quantizations
GGUF Q4
Fine-tune it
Permissive
QLoRA3.1 GB1× RTX 3060 12GB
LoRA5.7 GB1× RTX 3060 12GB
Full fine-tune30.8 GB1× RTX 5090 32GB
QLoRA SFT on ~10k samples ≈ $1.05 (1× RTX 3060 12GB)
H2O-Danube family
Elo progression across releases
API price weights · each benchmark row carries its own source badge (see methodology)