Benchmarks
Benchmarks
DeepSWE
Codingunit % · normalized over [0, 85]DeepSWE v1.1 — 113 long-horizon software-engineering tasks written from scratch to avoid contamination, measuring frontier coding agents on diverse real-world complexity. Mean over 5 trials.
#ModelSourceScoreNormalized
Score distribution
4 tracked results across the normalization window
085
Score vs. parameters
Open-weights models, log-x params
No open-weights models with disclosed parameter counts have a score here yet.