Benchmarks
Benchmarks

DeepSWE

Codingunit % · normalized over [0, 85]

DeepSWE v1.1 — 113 long-horizon software-engineering tasks written from scratch to avoid contamination, measuring frontier coding agents on diverse real-world complexity. Mean over 5 trials.

#ModelSourceScoreNormalized
Score distribution
4 tracked results across the normalization window
085
Score vs. parameters
Open-weights models, log-x params
No open-weights models with disclosed parameter counts have a score here yet.