Benchmarks
Benchmarks

CWE-Bench

Codingunit % · normalized over [0, 60]

CWE-Bench, run by Collinear — an external benchmark for automated vulnerability patching across CWE classes. Reported as pass@1.

#ModelSourceScoreNormalized
Score distribution
1 tracked result across the normalization window
060
Score vs. parameters
Open-weights models, log-x params
No open-weights models with disclosed parameter counts have a score here yet.