Benchmarks
Benchmarks
SWE-bench Multilingual
Codingunit % · normalized over [0, 92]SWE-bench Multilingual — 300 real-issue problems spanning 9 programming languages, extending SWE-bench beyond Python.
#ModelSourceScoreNormalized
Score distribution
5 tracked results across the normalization window
092
Score vs. parameters
Open-weights models, log-x params
Only one open-weights model with a disclosed parameter count — not enough to plot a trend.