Benchmarks
Benchmarks

SWE-Atlas Codebase QnA

Codingunit % · normalized over [0, 70]

SWE-Atlas Codebase QnA — 124 tasks over 11 real repositories testing deep code comprehension and question answering rather than patch generation. Mean pass@1 under a rubric judge.

#ModelSourceScoreNormalized
Score distribution
4 tracked results across the normalization window
070
Score vs. parameters
Open-weights models, log-x params
No open-weights models with disclosed parameter counts have a score here yet.