Benchmarks
Benchmarks
FrontierCode
Codingunit % · normalized over [0, 70]FrontierCode v1.1 (Main set) — 150 agentic software-engineering tasks derived from real open-source pull requests, run and graded by Cognition against held-out unit tests plus weighted code-quality rubrics. Composite score, mean@5.
#ModelSourceScoreNormalized
Score distribution
4 tracked results across the normalization window
070
Score vs. parameters
Open-weights models, log-x params
No open-weights models with disclosed parameter counts have a score here yet.