Benchmarks
Benchmarks
Terminal-Bench 3.0
Agentsunit % · normalized over [0, 45]Terminal-Bench 3.0 — the third-generation terminal-agent suite, a large step up in difficulty from 2.1 (frontier models score in the 20-35% range where they clear 85% on 2.1). General agent capability in a real shell, not coding alone.
#ModelSourceScoreNormalized
Score distribution
8 tracked results across the normalization window
045
Score vs. parameters
Open-weights models, log-x params