Benchmarks
Benchmarks

DeepSWE

Codingunit % · normalized over [0, 85]

DeepSWE v1.1 — 113 long-horizon software-engineering tasks written from scratch to avoid contamination, measuring frontier coding agents on diverse real-world complexity. Mean over 5 trials.

#ModelSourceScoreNormalized
Score distribution
16 tracked results across the normalization window
085
Score vs. parameters
Open-weights models, log-x params
1B10B100B1000B