Benchmarks
Benchmarks
FrontierBench
Agentsunit % · normalized over [0, 50]Frontier agent-work benchmark from the Terminal-Bench/Harbor team (formerly Terminal-Bench 3.0): 74 tasks across 7 domains in v0.1, spanning artifacts from databases and ML checkpoints to CAD files and formal proofs; resolution rate per agent-harness x model pair, with per-trial token and cost accounting.
#ModelSourceScoreNormalized
Score distribution
8 tracked results across the normalization window
050
Score vs. parameters
Open-weights models, log-x params
Only one open-weights model with a disclosed parameter count — not enough to plot a trend.