Benchmarks
Benchmarks
CWE-Bench
Codingunit % · normalized over [0, 60]CWE-Bench, run by Collinear — an external benchmark for automated vulnerability patching across CWE classes. Reported as pass@1.
#ModelSourceScoreNormalized
Score distribution
1 tracked result across the normalization window
060
Score vs. parameters
Open-weights models, log-x params
No open-weights models with disclosed parameter counts have a score here yet.