Benchmarks
Benchmarks

SWE-bench Multilingual

Codingunit % · normalized over [0, 92]

SWE-bench Multilingual — 300 real-issue problems spanning 9 programming languages, extending SWE-bench beyond Python.

#ModelSourceScoreNormalized
Score distribution
3 tracked results across the normalization window
092
Score vs. parameters
Open-weights models, log-x params
No open-weights models with disclosed parameter counts have a score here yet.