Benchmarks
Benchmarks

SWE-bench Multimodal

Codingunit % · normalized over [0, 70]

SWE-bench Multimodal — issues augmented with visual context (screenshots, design mockups), testing whether coding agents generalize to visual software domains.

#ModelSourceScoreNormalized
Score distribution
5 tracked results across the normalization window
070
Score vs. parameters
Open-weights models, log-x params
Only one open-weights model with a disclosed parameter count — not enough to plot a trend.