Benchmarks
Benchmarks
DeepSWE
Codingunit % · normalized over [0, 85]DeepSWE v1.1 — 113 long-horizon software-engineering tasks written from scratch to avoid contamination, measuring frontier coding agents on diverse real-world complexity. Mean over 5 trials.
#ModelSourceScoreNormalized
Score distribution
16 tracked results across the normalization window
085
Score vs. parameters
Open-weights models, log-x params